Senior NPU Compiler Engineer
PIMIC.ai · United States
Apply & track with Apply EdgeCompany Description:PIMIC.ai is an AI semiconductor and solutions company pioneering processing-in-memory hardware architectures for edge AI. This technology is designed to deliver up to 50x higher compute performance while reducing power consumption by up to 20x, enabling highly efficient AI inference across diverse applications. The company’s first product targets ultra-low-power speech recognition, achieving processing under 30 µA at roughly one-twentieth the power usage of leading current solutions. PIMIC.ai focuses on redefining how AI workloads run at the edge, creating opportunities for engineers to shape next-generation hardware–software co-designed systems. Team members collaborate in a fast-paced, innovation-driven environment that values technical excellence and practical impact.Role Description PIMIC is hiring a Senior NPU Compiler Engineer to build the software flow that imports neural-network models from TFLite, ONNX, and PyTorch-exported formats and compiles supported graphs into executable instruction and data images for PIMIC’s Processing-In-Memory NPU. The engineer will work on model import, supported-operator validation, shape inference, operator lowering, tensor layout conversion, memory planning, dataflow scheduling, instruction encoding, image generation, simulator integration, and model-output validation.Core qualifications:Strong C++ and Python programming skills.Experience with compiler, graph compiler, or code-generation tools.Experience with TFLite, ONNX, PyTorch FX / TorchScript / exported graphs, or similar graph IRs.Understanding of neural-network operators such as Conv1D/2D, DepthwiseConv, Dense, MatMul, Add, Multiply, Slice, Concat, GRU, SVDF, and activation functions.Experience with hardware-aware compilation, accelerator backends, DSP/NPU/GPU backends, or embedded inference runtimes.Familiarity with memory planning, tensor layout, scheduling, binary/image generation, and simulator-based validation.Ability to debug model execution differences between reference frameworks and accelerator/simulator output.Nice to have:MLIR, TVM, IREE, XLA, ONNX Runtime backend, or TFLite Micro experience.Prior work on NPU, DSP, GPU, FPGA, or ASIC accelerator compiler flows.Experience with instruction encoding, firmware image generation, or hardware simulator integration.Audio/voice model familiarity: KWS, DNR, ENC, SID, SVDF, GRU, compact CNNs.