Deep GPU programming and optimization experience (CUDA and/or Triton) - kernel-level tuning, memory hierarchy, and bandwidth optimization at scale. * Hands-on experience optimizing inference and ...
Deep GPU programming and optimization experience (CUDA and/or Triton) - kernel-level tuning, memory hierarchy, and bandwidth optimization at scale. * Hands-on experience optimizing inference and ...
Deep GPU programming and optimization experience (CUDA and/or Triton) - kernel-level tuning, memory hierarchy, and bandwidth optimization at scale. * Hands-on experience optimizing inference and ...
Deep GPU programming and optimization experience (CUDA and/or Triton) - kernel-level tuning, memory hierarchy, and bandwidth optimization at scale. * Hands-on experience optimizing inference and ...
Performance Engineer (Inference, Training & GPU)
San Francisco, CA · On-site
$200 - $300/hr
Deep GPU programming and optimization experience (CUDA and/or Triton) -- kernel-level tuning, memory hierarchy, and bandwidth optimization at scale. * Hands‑on experience optimizing inference and ...
Performance Engineer (Inference, Training & GPU)
San Francisco, CA · On-site
$200 - $300/hr
Deep GPU programming and optimization experience (CUDA and/or Triton) -- kernel-level tuning, memory hierarchy, and bandwidth optimization at scale. * Hands‑on experience optimizing inference and ...
Senior ML Accelerator Engineer - GPU
San Francisco, CA · On-site
$128.70 - $261.30/hr
Excellent GPU programming skills in CUDA, with a thorough understanding of parallel programming patterns and GPU architecture. * Hands‑on experience benchmarking, profiling, debugging and ...
Senior ML Accelerator Engineer - GPU
San Francisco, CA · On-site
$128.70 - $261.30/hr
Excellent GPU programming skills in CUDA, with a thorough understanding of parallel programming patterns and GPU architecture. * Hands‑on experience benchmarking, profiling, debugging and ...
Sr. AI Software Engineer (GPU/C++)
Milpitas, CA · On-site
$139K - $184K/yr
They are seeking a Sr. AI Software Engineer with a focus on C++ and GPU to design and implement core infrastructure components that support AI/ML workloads across various frameworks and hardware ...
Sr. AI Software Engineer (GPU/C++)
Milpitas, CA · On-site
$139K - $184K/yr
They are seeking a Sr. AI Software Engineer with a focus on C++ and GPU to design and implement core infrastructure components that support AI/ML workloads across various frameworks and hardware ...
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops
Santa Clara, CA · On-site
$150K/yr
Kernel engineering means demonstrating mastery in designing complex, scalable systems using modern C++, coupled with a fundamental grasp of GPU architectures (HIP/CUDA), memory hierarchies, and ...
Staff Software Development Engineer: GPU, Computer Vision, AI/ML Ops
Santa Clara, CA · On-site
$150K/yr
Kernel engineering means demonstrating mastery in designing complex, scalable systems using modern C++, coupled with a fundamental grasp of GPU architectures (HIP/CUDA), memory hierarchies, and ...
Senior Software Engineer, GPU Performance
Sunnyvale, CA · On-site
$143K - $189K/yr
Experience low-level GPU programming (CUDA, Triton, CUTLASS, etc.) and performance engineering techniques. * Experience with modern LLMs and their deployment on AI accelerators. Preferred ...
Senior Software Engineer, GPU Performance
Sunnyvale, CA · On-site
$143K - $189K/yr
Experience low-level GPU programming (CUDA, Triton, CUTLASS, etc.) and performance engineering techniques. * Experience with modern LLMs and their deployment on AI accelerators. Preferred ...
Proficiency in C , GPU Programming, and familiarity with Rust. About the job Google's software engineers develop the next-generation technologies that change how billions of users connect, explore ...
Proficiency in C , GPU Programming, and familiarity with Rust. About the job Google's software engineers develop the next-generation technologies that change how billions of users connect, explore ...
Senior ML Accelerator Engineer - GPU
San Francisco, CA · On-site
$128.70 - $261.30/hr
Excellent GPU programming skills in CUDA with a thorough understanding of parallel programming patterns and GPU architecture. * Hands‑on experience benchmarking, profiling, debugging, and ...
Senior ML Accelerator Engineer - GPU
San Francisco, CA · On-site
$128.70 - $261.30/hr
Excellent GPU programming skills in CUDA with a thorough understanding of parallel programming patterns and GPU architecture. * Hands‑on experience benchmarking, profiling, debugging, and ...
Kernel engineering means demonstrating mastery in designing complex, scalable systems using modern C++, coupled with a fundamental grasp of GPU architectures (HIP/CUDA), memory hierarchies, and ...
Kernel engineering means demonstrating mastery in designing complex, scalable systems using modern C++, coupled with a fundamental grasp of GPU architectures (HIP/CUDA), memory hierarchies, and ...
Senior Systems GPU Engineer - AI & Robotics
San Francisco, CA · On-site
$123K - $168K/yr
Primary Function of Position As a Senior Systems GPU Engineer - AI & Robotics, you will be responsible for iteratively improving and integrating high performance robotic AI with MLE product teams.
New
Senior Systems GPU Engineer - AI & Robotics
San Francisco, CA · On-site
$123K - $168K/yr
Primary Function of Position As a Senior Systems GPU Engineer - AI & Robotics, you will be responsible for iteratively improving and integrating high performance robotic AI with MLE product teams.
New
Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: As a leading technology innovator, Qualcomm pushes the boundaries of what's possible to enable next-generation ...
Engineering Group, Engineering Group > GPU ASICS Engineering General Summary: As a leading technology innovator, Qualcomm pushes the boundaries of what's possible to enable next-generation ...
Senior Systems GPU Engineer - AI & Robotics
$123K - $168K/yr
Primary Function of Position As a Senior Systems GPU Engineer - AI & Robotics, you will be responsible for iteratively improving and integrating high performance robotic AI with MLE product teams.
New
Senior Systems GPU Engineer - AI & Robotics
$123K - $168K/yr
Primary Function of Position As a Senior Systems GPU Engineer - AI & Robotics, you will be responsible for iteratively improving and integrating high performance robotic AI with MLE product teams.
New
Senior Systems GPU Engineer - AI & Robotics
San Francisco, CA · On-site
$123K - $168K/yr
Primary Function of Position As a Senior Systems GPU Engineer - AI & Robotics, you will be responsible for iteratively improving and integrating high performance robotic AI with MLE product teams.
New
Senior Systems GPU Engineer - AI & Robotics
San Francisco, CA · On-site
$123K - $168K/yr
Primary Function of Position As a Senior Systems GPU Engineer - AI & Robotics, you will be responsible for iteratively improving and integrating high performance robotic AI with MLE product teams.
New
We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits. The Role You'll be our performance ...
We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits. The Role You'll be our performance ...
Technical Program Manager - GPU AI
San Diego, CA · On-site
$171K - $256K/yr
Engineering Services Group, Engineering Services Group > Program Management General Summary: Qualcomm is seeking a Technical Program Manager to join the Adreno GPU Program Management team and drive ...
Technical Program Manager - GPU AI
San Diego, CA · On-site
$171K - $256K/yr
Engineering Services Group, Engineering Services Group > Program Management General Summary: Qualcomm is seeking a Technical Program Manager to join the Adreno GPU Program Management team and drive ...
GPU Software Engineer
San Jose, CA · On-site
GPU Software Engineer Location: San Jose, CA Duration: 6+ months contract (Long Term) Roles and Responsibilities: * As a GPU Software Engineer, you will be equipped to develop GPU IP from the early ...
GPU Software Engineer
San Jose, CA · On-site
GPU Software Engineer Location: San Jose, CA Duration: 6+ months contract (Long Term) Roles and Responsibilities: * As a GPU Software Engineer, you will be equipped to develop GPU IP from the early ...
We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits. The Role You'll be our performance ...
Quick apply
We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits. The Role You'll be our performance ...
Software Engineer, GPU
Mountain View, CA · On-site
$204K - $259K/yr
In this hybrid role, you will report to a Senior Software Engineer. You will: * Develop high-performance GPU primitives and abstractions to enable Waymo to scale its accelerator codebase across ...
Software Engineer, GPU
Mountain View, CA · On-site
$204K - $259K/yr
In this hybrid role, you will report to a Senior Software Engineer. You will: * Develop high-performance GPU primitives and abstractions to enable Waymo to scale its accelerator codebase across ...
GPU Image Processing Framework Software Engineer
$150K - $277K/yr
Description We are looking for a software engineer to join the Apple Camera and Photos team to help develop the next generation GPU-based imaging processing technologies on mobile and desktop ...
GPU Image Processing Framework Software Engineer
$150K - $277K/yr
Description We are looking for a software engineer to join the Apple Camera and Photos team to help develop the next generation GPU-based imaging processing technologies on mobile and desktop ...
Gpu Programmer information
See California salary details
$11.86 - $16.95
4% of jobs
$16.95 - $22.04
9% of jobs
$25.70 is the 25th percentile. Wages below this are outliers.
$22.04 - $27.13
17% of jobs
$27.13 - $32.22
13% of jobs
The median wage is $35.19 / hr.
$32.22 - $37.31
13% of jobs
$37.31 - $42.40
10% of jobs
$42.40 - $47.49
9% of jobs
$48.44 is the 75th percentile. Wages above this are outliers.
$47.49 - $52.58
9% of jobs
$52.58 - $57.67
7% of jobs
$57.67 - $62.76
6% of jobs
$62.76 - $67.85
4% of jobs
$11
$39
$67
How much do gpu programmer jobs pay per hour?
What are the key skills and qualifications needed to thrive in the GPU programmer position, and why are they important?
To thrive as a GPU Programmer, you need a solid background in computer science, experience with parallel computing concepts, and proficiency in GPU programming languages like CUDA or OpenCL. Familiarity with development tools such as NVIDIA Nsight, profiling utilities, and version control systems is typically required, while relevant certifications in GPU computing can be beneficial. Strong problem-solving ability, collaboration skills, and attention to detail help differentiate top performers in this field. These skills are essential for optimizing code performance, successfully working in dynamic teams, and meeting the high computational demands of modern applications.
What does a GPU programmer do?
A GPU Programmer specializes in writing and optimizing code that runs on Graphics Processing Units (GPUs). They use parallel computing techniques and languages like CUDA or OpenCL to accelerate tasks such as graphics rendering, scientific simulations, and machine learning. Their work involves optimizing performance, managing memory efficiently, and ensuring compatibility across different hardware architectures.

Full-time
Re-posted 10 days ago
Job description
We build foundational world models that can perceive, generate, reason, and interact with the 3D world - unlocking AI's full potential through spatial intelligence by transforming seeing into doing, perceiving into reasoning, and imagining into creating. We believe spatial intelligence will unlock new forms of storytelling, creativity, design, simulation, and immersive experiences across both virtual and physical worlds. We bring together a world-class team, united by a shared curiosity, passion, and deep backgrounds in technology - from AI research to systems engineering to product design - creating a tight feedback loop between our cutting-edge research and products that empower our users.
Role Overview
We are looking for a Performance Engineer to make World Labs' models train and serve as fast as the hardware allows.
Running large generative world models at scale is a novel systems problem. You will find the bottlenecks - in kernels, in the serving path, in the training loop, in how we use our GPUs - and eliminate them. Your ownership is technical and concrete: the throughput you unlock, the latency you cut, the utilization you win back, and the correctness you hold while doing it. You will work up and down the stack, from low-level tensor and kernel optimization to fleet-wide serving efficiency, in close partnership with the researchers whose models you are accelerating.
This is a hands-on, individual-contributor role. You will profile, design, build, and ship code directly.
What You Will Do:
- Optimize inference and serving end to end - latency, throughput, batching, caching, and scheduling - to serve our models efficiently at production scale.
- Write and tune GPU kernels (CUDA, Triton) for hot paths; drive kernel fusion, memory- and bandwidth-bound optimization, and low-precision (FP8/INT8) execution.
- Optimize training throughput and GPU utilization: parallelism strategies, communication/compute overlap, mixed precision, and eliminating pipeline stalls.
- Build performance models, profiling workflows, and observability that make throughput, latency, cost, utilization, and their tradeoffs legible across the stack.
- Own numerical correctness across precision, kernel, and hardware changes - treating correctness as part of performance, not separate from it.
- Partner with researchers to productionize models for serving and to make experiments run faster and more reliably.
- Where needed, work on the distributed systems that training and inference run on - but the core of the job is squeezing the most out of every GPU.
You should excel at the fundamentals below - we index on inference, serving, GPU optimization, and training performance. Distributed-systems breadth is welcome, but secondary.
- Strong performance-engineering foundations: profiling, roofline analysis, latency/throughput optimization, and disciplined root-cause investigation.
- Deep GPU programming and optimization experience (CUDA and/or Triton) - kernel-level tuning, memory hierarchy, and bandwidth optimization at scale.
- Hands-on experience optimizing inference and serving for large models: batching, KV/prompt caching, quantization, and low-latency, high-throughput sampling.
- Hands-on experience optimizing training performance: parallelism, distributed communication, mixed/low precision, and utilization.
- Working knowledge of ML framework internals (PyTorch and/or JAX; torch.compile, XLA, or similar compiler paths).
- Strong proficiency in Python, with the ability to drop into C++/CUDA (and Rust or Go) as the work demands.
- High-ownership mindset - you measure yourself by throughput shipped and latency cut, not tickets closed.
- Experience at an AI lab or ML-native company, optimizing systems used directly by researchers and productionizing research code.
- Low-precision and numerics depth: FP8/INT8 quantization, mixed-precision, and detecting numerical regressions across hardware platforms.
- Distributed systems for large-scale training and inference - collective communication (NCCL), interconnects (NVLink), model and tensor parallelism, and fault tolerance. A strong plus, but not a substitute for the core skills above.
- Experience serving generative, diffusion, video, or 3D/spatial models - not just text LLMs.
- Multi-accelerator experience (GPU plus TPU or Trainium) and partnering with hardware vendors on accelerator capabilities.
- Building performance-modeling and observability frameworks for GPU utilization and cost.
Who You Are:
- Fearless Innovator: We need people who thrive on challenges and aren't afraid to tackle the impossible.
- Resilient Builder: Impacting Large World Models isn't a sprint; it's a marathon with hurdles. We're looking for builders who can weather the storms of groundbreaking research and come out stronger.
- Mission-Driven Mindset: Everything we do is in service of creating the best spatially intelligent AI systems, and using them to empower people.
- Collaborative Spirit: We're building something bigger than any one person. We need team players who can harness the power of collective intelligence.
We're hiring the brightest minds from around the globe to bring diverse perspectives to our cutting-edge work. If you're ready to work on technology that will reshape how machines perceive and interact with the world, World Labs is your launchpad.
Join us, and let's make history together.
Equal Opportunity & Pay Transparency
Equal Employment Opportunity
World Labs is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, genetic information, veteran status, or any other characteristic protected under applicable law. We welcome all qualified applicants and are committed to providing reasonable accommodations throughout the hiring process upon request.
California Pay Transparency
In accordance with California law, we disclose the following:
Pay Range
$200-$300k base salary (good-faith estimate for San Francisco Bay Area upon hire; actual offer based on experience, skills, and qualifications)
Total Compensation
Base salary plus equity awards
Salary History
We do not request or consider prior compensation in making offers
Compliance: Cal. Lab. Code §432.3 (pay scale disclosure & salary history ban); Cal. Lab. Code §1197.5 (Equal Pay Act); Cal. Gov. Code §12940 (FEHA); 42 U.S.C. §2000e (Title VII); 29 U.S.C. §621 (ADEA); 42 U.S.C. §12101 (ADA)