1

Gpu Programmer Jobs in California (NOW HIRING)

GPU Kernel Engineer

San Francisco, CA · On-site

$190K - $250K/yr

About the role We are seeking a highly skilled GPU Kernel Engineer who is passionate about pushing the limits of performance on modern accelerators. In this role, you will design and optimize custom ...

Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: GPU System Driver Team are looking for talented software engineers to develop in-house GPU drivers to verify GPU function ...

Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: As a leading technology innovator, Qualcomm pushes the boundaries of what's possible to enable next-generation ...

Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: As a leading technology innovator, Qualcomm pushes the boundaries of what's possible to enable next-generation ...

GPU Engineer

Santa Clara, CA · On-site

$155K - $200K/yr

Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: This individual possesses the solid engineering fundamentals and understanding with some supervision in the architecting ...

As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models. You'll be responsible ...

Senior Compiler Engineer - Rust GPU

Santa Clara, CA · On-site

$143K - $189K/yr

We are redefining how developers write high-performance GPU software by bringing the safety, expressiveness, and modern tooling of Rust to native GPU and CUDA development. On this team, you will ...

Senior Compiler Engineer - Rust GPU

Santa Clara, CA · On-site

$143K - $189K/yr

We are redefining how developers write high-performance GPU software by bringing the safety, expressiveness, and modern tooling of Rust to native GPU and CUDA development. On this team, you will ...

Showing results 21-40

Gpu Programmer information

See California salary details

$11

$39

$67

How much do gpu programmer jobs pay per hour?

As of Sep 3, 2026, the average hourly pay for gpu programmer in California is $39.02, according to ZipRecruiter salary data. Most workers in this role earn between $25.38 and $50.77 per hour, depending on experience, location, and employer.

What does a GPU programmer do?

A GPU Programmer specializes in writing and optimizing code that runs on Graphics Processing Units (GPUs). They use parallel computing techniques and languages like CUDA or OpenCL to accelerate tasks such as graphics rendering, scientific simulations, and machine learning. Their work involves optimizing performance, managing memory efficiently, and ensuring compatibility across different hardware architectures.

What are the key skills and qualifications needed to thrive in the GPU programmer position, and why are they important?

To thrive as a GPU Programmer, you need a solid background in computer science, experience with parallel computing concepts, and proficiency in GPU programming languages like CUDA or OpenCL. Familiarity with development tools such as NVIDIA Nsight, profiling utilities, and version control systems is typically required, while relevant certifications in GPU computing can be beneficial. Strong problem-solving ability, collaboration skills, and attention to detail help differentiate top performers in this field. These skills are essential for optimizing code performance, successfully working in dynamic teams, and meeting the high computational demands of modern applications.

What are popular job titles related to Gpu Programmer jobs in California?

For Gpu Programmer jobs in California, the most frequently searched job titles are:

What job categories do people searching Gpu Programmer jobs in California look for?

The top searched job categories for Gpu Programmer jobs in California are:

What cities in California are hiring for Gpu Programmer jobs?

Cities in California with the most Gpu Programmer job openings:

Infographic showing various Gpu Programmer job openings in California as of August 2026, with employment types broken down into 87% Full Time, 6% Part Time, 6% Contract, and 1% Nights. Highlights an 89% Physical, 3% Hybrid, and 8% Remote job distribution, with an average salary of $81,158 per year, or $39 per hour.

Performance Engineer, GPU

Anthropic

San Francisco, CA

Full-time

Re-posted 15 days ago


Job description

About the role:

Pioneering the next generation of AI requires breakthrough innovations in GPU performance and systems engineering. As a GPU Performance Engineer, you'll architect and implement the foundational systems that power Claude and push the frontiers of what's possible with large language models. You'll be responsible for maximizing GPU utilization and performance at unprecedented scale, developing cutting-edge optimizations that directly enable new model capabilities and dramatically improve inference efficiency.

Working at the intersection of hardware and software, you'll implement state-of-the-art techniques from custom kernel development to distributed system architectures. Your work will span the entire stack-from low-level tensor core optimizations to orchestrating thousands of GPUs in perfect synchronization.

Strong candidates will have a track record of delivering transformative GPU performance improvements in production ML systems and will be excited to shape the future of AI infrastructure alongside world-class researchers and engineers.

You might be a good fit if you:
  • Have deep experience with GPU programming and optimization at scale
  • Are impact-driven, passionate about delivering measurable performance breakthroughs
  • Can navigate complex systems from hardware interfaces to high-level ML frameworks
  • Enjoy collaborative problem-solving and pair programming
  • Want to work on state-of-the-art language models with real-world impact
  • Care about the societal impacts of your work
  • Thrive in ambiguous environments where you define the path forward
Strong candidates may also have experience with:
  • GPU Kernel Development: CUDA, Triton, CUTLASS, Flash Attention, tensor core optimization
  • ML Compilers & Frameworks: PyTorch/JAX internals, torch.compile, XLA, custom operators
  • Performance Engineering: Kernel fusion, memory bandwidth optimization, profiling with Nsight
  • Distributed Systems: NCCL, NVLink, collective communication, model parallelism
  • Low-Precision: INT8/FP8 quantization, mixed-precision techniques
  • Production Systems: Large-scale training infrastructure, fault tolerance, cluster orchestration
Representative projects:
  • Co-design attention mechanisms and algorithms for next-generation hardware architectures
  • Develop custom kernels for emerging quantization formats and mixed-precision techniques
  • Design distributed communication strategies for multi-node GPU clusters
  • Optimize end-to-end training and inference pipelines for frontier language models
  • Build performance modeling frameworks to predict and optimize GPU utilization
  • Implement kernel fusion strategies to minimize memory bandwidth bottlenecks
  • Create resilient systems for planet-scale distributed training infrastructure
  • Profile and eliminate performance bottlenecks in production serving infrastructure
  • Partner with hardware vendors to influence future accelerator capabilities and software stacks
 

Deadline to apply: None. Applications will be reviewed on a rolling basis. 

The expected salary range for this position is: