1

Npu Compiler Jobs (NOW HIRING)

Build deep fluency across the Chimera stack - quantization, custom kernels, and compiler tooling ... Background at NPU/edge-AI IP and silicon companies (e.g., Qualcomm SNPE/QNN, NVIDIA TensorRT/Jetson ...

Experience with NPU, DSP, or GPU-class accelerators * Experience running or debugging ML inference workloads on pre-silicon platforms * Knowledge of AI software stacks (runtime, compiler, graph ...

$175K - $225K/yr

Build deep fluency across the Chimera stack -- quantization, custom kernels, and compiler tooling ... Background at NPU/edge-AI IP and silicon companies (e.g., Qualcomm SNPE/QNN, NVIDIA TensorRT/Jetson ...

Forward Deployed Engineer

Burlingame, CA · On-site

$175K - $225K/yr

Build deep fluency across the Chimera stack -- quantization, custom kernels, and compiler tooling ... Background at NPU/edge-AI IP and silicon companies (e.g., Qualcomm SNPE/QNN, NVIDIA TensorRT/Jetson ...

Systems Performance Architect

Cupertino, CA · On-site

$206K/yr

... CPU/GPU/NPU microarchitecture concepts (pipelines, caches, coherence, etc.) Experience in ... Expertise in one or more focused areas (CPU/GPU, System Architecture, Compiler/SW Stack). Areas ...

AI Infrastructure Engineer

San Jose, CA · On-site

$126K - $165K/yr

Proficiency with GPU/NPU programming (CUDA, or vendor-specific SDKs), compiler toolchains, and deep learning frameworks (PyTorch, or TensorFlow). * Strong programming skills in C/C++, with a track ...

AI Infrastructure Engineer

San Jose, CA

$126K - $165K/yr

Proficiency with GPU/NPU programming (CUDA, or vendor-specific SDKs), compiler toolchains, and deep learning frameworks (PyTorch, or TensorFlow). * Strong programming skills in C/C++, with a track ...

Interacting with OS, Compiler, and ML software teams to understand future software trends and guide ... CPU/GPU/NPU microarchitecture concepts (pipelines, caches, coherence, etc.) Experience in ...

Proficiency with GPU/NPU programming (CUDA, or vendor-specific SDKs), compiler toolchains, and deep learning frameworks (PyTorch, or TensorFlow). * Strong programming skills in C/C++, with a track ...

Interacting with OS, Compiler, and ML software teams to understand future software trends and guide ... Understanding of CPU/GPU/NPU microarchitecture concepts (pipelines, caches, coherence, etc.

Showing results 41-60

Npu Compiler information

What is an NPU compiler?

An NPU Compiler is specialized software that translates high-level machine learning models or code into instructions optimized for Neural Processing Units (NPUs). NPUs are hardware accelerators designed to efficiently execute deep learning and AI workloads. The compiler bridges the gap between standard AI frameworks and the unique architecture of NPUs, optimizing computations for speed and power efficiency. It ensures that neural network operations are executed correctly and efficiently, often handling tasks like operator fusion, quantization, and hardware-specific optimization.

What are the key skills and qualifications needed to thrive as an NPU compiler engineer?

To thrive as an NPU Compiler Engineer, you need a strong background in computer science, compiler theory, and experience with neural network architectures, often supported by a degree in computer engineering or a related field. Familiarity with programming languages like C++, Python, and frameworks such as TensorFlow or PyTorch, as well as experience with hardware description languages and specialized compilers, is typically required. Excellent problem-solving skills, attention to detail, and effective collaboration are vital soft skills in this role. These competencies are crucial for optimizing neural network models to run efficiently on specialized hardware, ensuring high-performance and reliable AI solutions.

What are some common challenges faced by NPU compiler engineers, and how can they be addressed?

NPU Compiler engineers often encounter challenges related to optimizing code for specialized hardware, such as balancing performance with power consumption and ensuring compatibility across different neural processing unit architectures. Debugging and profiling code can also be complex due to the parallel nature of NPUs and limited visibility into hardware operations. To address these challenges, engineers frequently collaborate with hardware teams, utilize advanced profiling tools, and stay updated on the latest compiler optimization techniques. Building a strong foundation in both software and hardware principles is essential for success in this role.

What is the difference between Npu Compiler vs Hardware Engineer?

AspectNpu CompilerHardware Engineer
Required CredentialsBachelor's or higher in Computer Science, Electrical Engineering, or related fields; knowledge of hardware description languagesBachelor's or higher in Electrical Engineering, Computer Engineering, or related fields; certifications like CCNA or Cisco certifications are common
Work EnvironmentSoftware development teams, AI and hardware integration labs, R&D departmentsDesign labs, manufacturing facilities, R&D departments
Industry UsageAI hardware, embedded systems, chip design companiesSemiconductor companies, consumer electronics, hardware manufacturing

The Npu Compiler focuses on developing software tools that optimize and translate neural network models for AI hardware, while Hardware Engineers design and develop physical hardware components. Both roles often collaborate but serve different parts of the hardware-software ecosystem in AI and electronics industries.

What are popular job titles related to Npu Compiler jobs?

For Npu Compiler jobs, the most frequently searched job titles are:

Infographic showing various Npu Compiler job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 95% Full Time, 1% Part Time, 1% Temporary, and 2% Contract. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution.

Senior/Staff Software Engineer - Embedded AI Runtime

San Diego, CA • On-site

$131K - $172K/yr

Other

Posted 19 days ago


Job description

Senior/Staff Software Engineer – Embedded AI Runtime Senior/Staff Software Engineer – Embedded AI Runtime

Location: San Diego, California
Focus: C/C++, Embedded AI, AI Runtimes, IoT

TalentLab is recruiting a Senior/Staff Software Engineer to join a team developing high-performance AI runtime software for IoT platforms.

This is a hands-on software engineering role focused on the infrastructure that enables modern AI models to execute efficiently on-device. Rather than building or training models, you'll work deeper in the AI software stack — developing runtime and SDK capabilities, integrating with edge and IoT platforms, and solving performance and compatibility challenges across hardware and software.

You'll work with advanced neural network architectures including DNNs and LLMs while collaborating closely with hardware, operating systems, compiler, driver and AI software teams.

What You'll Work On

Develop and enhance AI runtime and SDK capabilities for embedded and IoT platforms.

Lead significant features from technical design through implementation, debugging and integration.

Develop high-performance systems software primarily in C/C++.

Enable efficient execution of DNNs, LLMs and other modern neural network architectures.

Integrate AI software across multiple embedded product platforms.

Debug complex issues spanning AI runtime software, operating systems, compilers, drivers and hardware.

Work closely with platform, hardware and software engineering teams.

Evaluate new developments in AI and systems software and determine how they can improve the runtime platform.

Provide technical guidance to other engineers and contribute to broader architecture and technical planning.

What We're Looking For

Strong professional software development experience, ideally 6+ years, with deeper experience expected at the Staff level.

Advanced C/C++ development skills.

Experience developing systems software in Linux or Unix environments.

Strong software engineering fundamentals including data structures, algorithms, object-oriented design, debugging and testing.

Experience building embedded software or working close to hardware.

Ability to independently own complex technical problems and drive them through to completion.

Experience providing technical leadership or guidance within a software engineering team.

Bachelor's, Master's or PhD in Computer Science, Computer Engineering, Electrical Engineering or a related discipline.

Particularly Relevant Experience

We're especially interested in candidates with experience in one or more of the following:

AI runtimes, inference engines or AI SDK development.

Embedded or on-device AI.

DNN, CNN, RNN/LSTM, LLM or other neural network architectures.

Hardware-accelerated AI inference.

Low-level interactions between operating systems and hardware.

Linux, Android, QNX or similar embedded operating systems.

Debugging across hardware, OS, compiler and driver layers.

AI accelerator, DSP, GPU or NPU software.

Agile development and Git-based source control.

This is a strong fit for an experienced systems software engineer who wants to work underneath the models, solving the runtime, performance and platform challenges required to bring advanced AI capabilities onto real-world embedded devices.

#J-18808-Ljbffr