Senior NPU Compiler Engineer
Build the compiler stack for PIMIC’s ultra-low-power Processing-in-Memory NPU
Employment Type: Full-time
About PIMIC
PIMIC is developing ultra-low-power neural-processing technology for edge-AI applications. Our Processing-in-Memory architecture combines high computational efficiency with a dataflow-based execution model, enabling advanced AI workloads within stringent power and memory constraints.
We are seeking a Senior NPU Compiler Engineer to develop the compiler technology that maps machine-learning models onto PIMIC’s proprietary NPU architecture.
Position Summary
The Senior NPU Compiler Engineer will design and implement a compiler that imports models from frameworks and formats such as TFLite, ONNX, and PyTorch-exported graphs and converts supported neural-network operations into executable instruction and data images for PIMIC’s Processing-in-Memory NPU.
The role covers the complete compilation flow, including graph import, operator validation, shape inference, graph transformations, operator lowering, tensor-layout selection, memory planning, scheduling, instruction encoding, image generation, simulator integration, and output validation.
The engineer will work closely with the NPU architecture, RTL, simulator, firmware, and machine-learning teams to define the software-to-hardware interface and deliver a reliable, production-quality compiler toolchain.
Responsibilities
Design and develop the compiler toolchain for PIMIC’s proprietary NPU architecture.
Import and process machine-learning models from TFLite, ONNX, and PyTorch-exported formats.
Parse and validate computational graphs, operators, tensors, shapes, attributes, and dependencies.
Implement shape inference, constant propagation, graph transformations, and operator fusion where applicable.
Lower high-level neural-network operators into PIMIC NPU primitives and instruction sequences.
Map neural-network graphs onto PIMIC’s dataflow execution architecture.
Develop scheduling algorithms for operators, data transfers, and hardware resources.
Implement tensor-layout transformations and optimize data placement for the target architecture.
Develop memory-planning and allocation algorithms for weights, activations, intermediate data, and instruction storage.
Generate executable instruction images and associated data images for NPU memories.
Define and maintain compiler interfaces with the NPU simulator, firmware, runtime, and RTL environment.
Develop validation infrastructure to compare compiler-generated results with framework reference outputs.
Diagnose functional and performance differences across the framework, compiler, simulator, and RTL implementations.
Optimize generated programs for execution efficiency, memory utilization, bandwidth, and power.
Create debugging, visualization, tracing, and performance-analysis tools for compiled models.
Define compiler requirements and hardware/software interfaces in collaboration with NPU architects.
Support new operators, architectural features, and customer models as the NPU platform evolves.
Develop automated unit, regression, and model-level tests for the compiler toolchain.
Produce clear design documentation and maintain high-quality, production-ready software.
Required Qualifications
Bachelor’s or Master’s degree in Computer Science, Computer Engineering, Electrical Engineering, or a related field.
Significant experience developing compilers, graph compilers, code generators, or backend toolchains.
Strong programming skills in C++ and Python.
Solid understanding of compiler concepts, including intermediate representations, graph transformations, lowering, scheduling, memory allocation, and code generation.
Experience working with computational graphs or machine-learning model formats such as ONNX, TFLite, or PyTorch-exported graphs.
Understanding of neural-network operations, tensor shapes, tensor layouts, and execution dependencies.
Experience developing software for specialized processors, accelerators, DSPs, GPUs, NPUs, or embedded systems.
Strong knowledge of computer architecture, memory systems, data movement, and performance optimization.
Experience developing automated test and validation infrastructure.
Strong debugging and problem-solving skills across multiple levels of the software and hardware stack.
Ability to work effectively with architecture, RTL, firmware, simulator, and machine-learning teams.
Preferred Qualifications
Experience with MLIR, LLVM, TVM, XLA, Glow, IREE, or another compiler infrastructure.
Experience developing a compiler backend for an NPU, DSP, GPU, or other domain-specific accelerator.
Familiarity with dataflow architectures, Processing-in-Memory architectures, or systolic accelerators.
Experience with operator fusion, tiling, scheduling, tensor-layout optimization, and scratchpad-memory management.
Experience generating binary instruction streams, memory images, or firmware-consumable artifacts.
Familiarity with DMA engines, on-chip SRAM, external memory, and memory-bandwidth optimization.
Experience correlating results among machine-learning frameworks, instruction-set simulators, cycle-accurate models, FPGA prototypes, and RTL.
Familiarity with embedded-AI deployment and resource-constrained edge devices.
Experience building compiler diagnostics, graph visualization, profiling, or performance-analysis tools.
What We Offer
An opportunity to build a compiler and software stack for a new ultra-low-power NPU architecture.
Direct involvement in architectural and hardware/software co-design decisions.
Work on production-oriented edge-AI technology from model import through silicon execution.
A collaborative environment with significant technical ownership and impact.
A competitive compensation package including stock options
Location
India or USA. Position could be located anywhere in these two countries (remote), or in-person in Chennai or the Bay Area