1

Neural Processing Unit Engineer Jobs (NOW HIRING)

AI Performance Modeling Engineer

Burlingame, CA ยท On-site

$180K - $225K/yr

About Quadric Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C ...

Full-Stack Software Engineer

Burlingame, CA ยท On-site

$110K - $270K/yr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Role The Full-Stack Engineer is key to making the Quadric product and toolchain easily accessible ...

Full-Stack Software Engineer

Burlingame, CA ยท On-site

$110K - $270K/yr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Role The Full-Stack Engineer is key to making the Quadric product and toolchain easily accessible ...

Deep Learning Compiler Engineer

Burlingame, CA ยท Remote

$110K - $270K/yr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Role As a senior member of our platform software engineering team, you will be tasked with lowering ...

Full-Stack Software Engineer

Burlingame, CA ยท On-site

$110K - $270K/yr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Role The Full-Stack Engineer is key to making the Quadric product and toolchain easily accessible ...

AI Inference Engineer

Burlingame, CA ยท On-site

$110K - $270K/yr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Role The AI Inference Engineer in Quadric is the key bridge between the world of AI/LLM models and ...

Deep Learning Compiler Engineer

Burlingame, CA ยท On-site +1

$110K - $270K/yr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Role As a senior member of our platform software engineering team, you will be tasked with lowering ...

Design Verification Intern

Burlingame, CA ยท On-site

$45 - $60/hr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Currently pursuing a Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, or a ...

Design Verification Intern

Burlingame, CA ยท On-site

$45 - $60/hr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Currently pursuing a Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, or a ...

Forward Deployed Engineer

Burlingame, CA ยท On-site

$175K - $225K/yr

About Quadric Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C ...

Design Verification Intern

Burlingame, CA ยท On-site

$45 - $60/hr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Currently pursuing a Bachelor's, Master's, or PhD in Computer Science, Electrical Engineering, or a ...

AI Kernel Engineer

Burlingame, CA ยท On-site

$110K - $270K/yr

Quadric has created an innovative general purpose neural processing unit (GPNPU) architecture ... Role The AI Kernel Engineer in Quadric plays the key role to enable a large number of AI kernels ...

$175K - $225K/yr

About Quadric Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C ...

Staff SoC RTL Engineer

Burlingame, CA ยท On-site

$175K - $230K/yr

About Quadric Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C ...

next page

Showing results 1-20

Neural Processing Unit Engineer information

See salary details

$49.5K

$113.5K

How much do neural processing unit engineer jobs pay per year?

As of Sep 10, 2026, the average yearly pay for neural processing unit engineer in the United States is $108,847.00, according to ZipRecruiter salary data. Most workers in this role earn between $113,000.00 and $113,000.00 per year, depending on experience, location, and employer.

What is a neural processing unit engineer?

A Neural Processing Unit (NPU) Engineer is a specialized hardware or software engineer who designs, develops, and optimizes processors specifically built to accelerate artificial intelligence (AI) and machine learning tasks. These engineers work on creating efficient NPU architectures, writing low-level code, and integrating NPUs into various computing systems such as smartphones, edge devices, and data centers. Their main goal is to maximize the performance and energy efficiency of AI workloads, such as neural network inference and training, by leveraging dedicated hardware. NPU Engineers often collaborate with data scientists, software developers, and hardware teams to ensure seamless deployment of AI models.

What are the key skills and qualifications needed to thrive as a neural processing unit engineer?

To thrive as a Neural Processing Unit Engineer, you need a solid background in computer engineering, digital circuit design, and deep learning algorithms, often supported by a relevant degree in electrical engineering or computer science. Familiarity with hardware description languages (HDL), simulation tools like ModelSim, and frameworks such as TensorFlow or PyTorch is typically required. Strong problem-solving skills, attention to detail, and effective teamwork set top performers apart in this role. These competencies are crucial for designing efficient NPUs that accelerate AI workloads and meet evolving performance and energy efficiency demands.

What are some common challenges faced by neural processing unit engineers when optimizing hardware for AI workloads?

Neural Processing Unit (NPU) Engineers often encounter challenges in balancing performance, power efficiency, and scalability while designing hardware for AI workloads. Achieving low latency and high throughput for diverse neural network models requires innovative architecture and close collaboration with software teams. Additionally, NPUs must be flexible enough to support emerging AI algorithms, which means engineers need to stay current with rapid advancements in machine learning. Overcoming these challenges typically involves extensive simulation, benchmarking, and iterative hardware-software co-optimization.

What is the difference between Neural Processing Unit Engineer vs AI Hardware Engineer?

AspectNeural Processing Unit EngineerAI Hardware Engineer
CredentialsBachelor's or Master's in Electrical Engineering, Computer Engineering, or related fields; experience with hardware design and AI acceleratorsBachelor's or Master's in Electrical Engineering, Computer Engineering, or related fields; focus on hardware development for AI systems
Work EnvironmentDesigning, testing, and optimizing neural processing units in R&D labs or tech companiesDeveloping and integrating AI hardware components in product development or research settings
Industry UsagePrimarily in AI chip design, machine learning hardware accelerationBroader AI hardware development including processors, accelerators, and embedded systems

Neural Processing Unit Engineers focus specifically on designing and optimizing neural processing units for AI applications, while AI Hardware Engineers work on a wider range of AI hardware components. Both roles require similar technical backgrounds but differ in scope and specialization within AI hardware development.

What are popular job titles related to Neural Processing Unit Engineer jobs?

For Neural Processing Unit Engineer jobs, the most frequently searched job titles are:

Infographic showing various Neural Processing Unit Engineer job openings in the United States as of September 2026, with employment types broken down into 2% As Needed, 80% Full Time, 11% Part Time, and 7% Contract. Highlights an 95% Physical, 1% Hybrid, and 4% Remote job distribution, with an average salary of $108,847 per year, or $52.3 per hour.

AI Performance Modeling Engineer

Burlingame, CA โ€ข On-site

quadric, Inc
Semiconductor and Electronic Component Manufacturingย โ€ขย 11 - 50 employees

$180K - $225K/yr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

This job post hasย expired today.ย Applications are no longer accepted.


Job description

About Quadric

Quadric is redefining edge AI with the industry's first General Purpose Neural Processing Unit (GPNPU), enabling developers to run both neural network inference and conventional C++ code on a single programmable architecture. Our technology powers intelligent edge devices across automotive, industrial, robotics, and embedded systems.

Founded by technologists from MIT and Carnegie Mellon, Quadric is a well-funded growth-stage semiconductor IP company with a growing licensing business.

The Opportunity

Quadric has created an innovative General-Purpose Neural Processing Unit (GPNPU) architecture. Unlike standard accelerators, the Quadric GPNPU executes both neural network graph code and conventional C++ DSP/control code across edge and endpoint devices.

As an AI Performance Modeling Engineer, you will build analytical, cycle-level performance models of AI inference workloads on our next-generation architecture in Python before silicon exists. These models directly guide team decisions on hardware lane bindings, tensor placement, and architecture trade-offs. We welcome candidates across all experience levels—from early-career engineers to seasoned experts—with direct mentorship provided to help you master mapping complex workloads (like LLMs) onto our custom hardware.

What You'll DoPerformance Modeling & Architectural Analysis
  • Build analytical, cycle-level Python models of AI inference workloads executing on next-generation GPNPU hardware.
  • Derive from first principles which hardware lanes operations bind on (compute, on-chip/external memory bandwidth, interconnect) and model software pipelining overlaps.
  • Model tensor placement, tiling across processing elements, local memory residency, and data movement across memory tiers.
  • Model sharding and collective boundary communication across multi-die systems.
Workload Adaptation & Technical Writing
  • Incorporate architectural details across vision networks and Large Language Models (LLMs), including operator mix, sparsity, routing, and quantization/low-precision numeric formats.
  • Calibrate performance models against an instruction-set simulator and profiling traces to meet stated accuracy targets.
  • Write and defend technical studies presenting empirical evidence that directly informs architecture and product decisions.
  • Balance single-stream latency against scaled throughput performance.
What Success Looks Like

Within your first 6–12 months, you'll:

  • Own a full workload's model end to end, calibrated against simulation and trusted by the engineering team.
  • Build performance models that consistently predict workload behavior within 20–30% of actual measurements.
  • Publish a written study whose defended conclusions directly shape an architecture or product decision.
  • Review and extend performance models beyond your initial starting domain.
What We're Looking ForRequired
  • Python & Quantitative Modeling: Strong Python skills with experience writing, validating, and calibrating numerical or quantitative models in code.
  • Computer Architecture Fundamentals: Solid grasp of memory hierarchies, bandwidth/latency trade-offs, pipelining, and execution bottlenecks (via industry experience, coursework, or research).
  • Technical Writing: Comfort writing clear technical studies that state and defend evidence-based conclusions.
  • Core Technical Depth (One of the following):
    • Option A: Deep understanding of NN inference operators and tensor shapes (e.g., Transformers, attention mechanisms, MoE, prefill/decode split).
    • Option B: Proven performance modeling experience in another quantitative/technical domain.
  • Education: BS, MS, or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or equivalent practical experience.
Preferred
  • Prior experience with GPUs, custom AI accelerators, CUDA, or Triton kernels.
  • Familiarity with roofline analysis, back-of-the-envelope estimation, or architecture simulators (e.g., gem5, Timeloop, MAESTRO, Accel-Sim).
  • Background in compiler internals (cost models, autotuners) or proficiency in C++.
  • Published performance studies or technical write-ups.
What We Offer

The base salary range for this position is $180,000 to $225,000. This range reflects the full span of levels and geographies at which Quadric hires for this role. The actual base salary offered will depend on a number of factors, including the specific level of the role, years and depth of relevant experience, technical skills and competencies, the criticality of the role to the business, internal equity, and work location. In addition to base salary, this role is eligible for equity and a discretionary annual performance bonus as applicable to the role and level. 

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings; for roles in other locations, benefits vary and are shared during the hiring process. These include:

  • Medical, dental, and vision insurance from day one - Premiums covered at 99% for Employees
  • Company-paid life Insurance 
  • Voluntary supplemental life insurance 
  • STD + LTD insurance 
  • Commuter support including parking or Caltrain reimbursement. Our office is conveniently located within walking distance of the Caltrain station
  • FSA + HSA
  • Equity with the business
  • Paid Parental Leave 
  • 401(k) Retirement Plan
  • Flexible PTO
  • Winter holiday shutdown
  • Catered lunch each day in our office 
  • Downtown Burlingame office location, close to shops, cafes, and local amenities
  • Collaborative, low-ego culture with significant ownership and impact
  • A work culture focused on innovative disruption

Founded in 2016 and based in downtown Burlingame, California, Quadric is building the world’s first supercomputer designed for the real-time needs of edge devices. Quadric aims to empower developers in every industry with superpowers to create tomorrow’s technology, today. The company was co-founded by technologists from MIT and Carnegie Mellon, who were previously the technical co-founders of the Bitcoin computing company 21.

Quadric is proud to be an equal opportunity employer. We are committed to creating an inclusive environment where people from all backgrounds can do their best work. We consider all qualified applicants without regard to race, color, religion, sex, gender identity or expression, sexual orientation, national origin, age, disability, veteran status, or any other protected characteristic under applicable law.

If this role resonates with you, we encourage you to apply even if your experience does not perfectly match every qualification. We value potential, curiosity, and a willingness to learn just as much as direct experience. Skills and growth come in many forms, and we would love to hear your story.

By submitting an application, you acknowledge that Quadric will collect and process your personal information as part of the hiring process. Please review our Privacy Policy to understand how we handle your data.


Quadric.io logo

About Quadric.io

Sourced by ZipRecruiter

Industry

Semiconductor and electronic component manufacturing

Company size

11 - 50 Employees

Headquarters location

Burlingame, CA, US

Year founded

2016

Social media