CUDA features improve both productivity and performance of AI applications. Your work in AI toolkits will accelerate enabling those for the community. This is an outstanding opportunity for someone ...
CUDA features improve both productivity and performance of AI applications. Your work in AI toolkits will accelerate enabling those for the community. This is an outstanding opportunity for someone ...
We are looking for a seasoned software professional to work on the CUDA Driver, a core component of our platform for accelerating general purpose computation on the GPU. You will be an integral part ...
We are looking for a seasoned software professional to work on the CUDA Driver, a core component of our platform for accelerating general purpose computation on the GPU. You will be an integral part ...
OR · On-site
$122K - $161K/yr
CUDA Tile shipped with CUDA 13.1 and is a major addition to CUDA ( You will design and implement compiler transformations, develop MLIR-based dialects and lowering passes, and optimize the ...
Senior Software Engineer, CUDA Deep Learning Systems
$121K - $160K/yr
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Senior Software Engineer, CUDA Deep Learning Systems
$121K - $160K/yr
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Senior Software Engineer - CUDA and Unified Memory
Santa Clara, CA · On-site
$142K - $188K/yr
CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, and ...
Senior Software Engineer - CUDA and Unified Memory
Santa Clara, CA · On-site
$142K - $188K/yr
CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, and ...
Senior Deep Learning Frameworks CUDA Software Engineer
Santa Clara, CA · On-site
$143K - $189K/yr
CUDA features improve both productivity and performance of AI applications. Your work in AI toolkits will accelerate enabling those for the community. This is an outstanding opportunity for someone ...
Senior Deep Learning Frameworks CUDA Software Engineer
Santa Clara, CA · On-site
$143K - $189K/yr
CUDA features improve both productivity and performance of AI applications. Your work in AI toolkits will accelerate enabling those for the community. This is an outstanding opportunity for someone ...
Dalcom Engineering is currently seeking a software developer skilled in C++ and CUDA programming, to support Science and Technology (S&T) efforts for radar and signal systems at Aberdeen Proving ...
Dalcom Engineering is currently seeking a software developer skilled in C++ and CUDA programming, to support Science and Technology (S&T) efforts for radar and signal systems at Aberdeen Proving ...
Senior DL Compiler Engineer -CUDA Tile
Austin, TX · On-site
$121K - $160K/yr
CUDA Tile shipped with CUDA 13.1 and is a major addition to CUDA ( You will design and implement compiler transformations, develop MLIR-based dialects and lowering passes, and optimize the ...
Senior DL Compiler Engineer -CUDA Tile
Austin, TX · On-site
$121K - $160K/yr
CUDA Tile shipped with CUDA 13.1 and is a major addition to CUDA ( You will design and implement compiler transformations, develop MLIR-based dialects and lowering passes, and optimize the ...
Senior DL Compiler Engineer -CUDA Tile
Redmond, WA · On-site
$137K - $180K/yr
CUDA Tile shipped with CUDA 13.1 and is a major addition to CUDA ( You will design and implement compiler transformations, develop MLIR-based dialects and lowering passes, and optimize the ...
Senior DL Compiler Engineer -CUDA Tile
Redmond, WA · On-site
$137K - $180K/yr
CUDA Tile shipped with CUDA 13.1 and is a major addition to CUDA ( You will design and implement compiler transformations, develop MLIR-based dialects and lowering passes, and optimize the ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
Our team is dedicated to exploring and prototyping next-generation ideas that bridge the gap between deep learning algorithms and CUDA, pushing the boundaries of what is possible on modern ...
NVIDIA's core CUDA math libraries are the foundation of NVIDIA CUDA-X libraries. This position will require leading at the intersection of Engineering, Product, and our customers to advocate ...
NVIDIA's core CUDA math libraries are the foundation of NVIDIA CUDA-X libraries. This position will require leading at the intersection of Engineering, Product, and our customers to advocate ...
At the center of this platform are CUDA Core Libraries that provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software. We ...
At the center of this platform are CUDA Core Libraries that provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software. We ...
At the center of this platform are CUDA Core Libraries that provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software. We ...
At the center of this platform are CUDA Core Libraries that provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software. We ...
At the center of this platform are CUDA Core Libraries that provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software. We ...
At the center of this platform are CUDA Core Libraries that provide the algorithms, abstractions, and runtime capabilities needed to build fast, reliable, and scalable GPU-accelerated software. We ...
Senior Software Engineer, CUDA Python Core Libraries
Santa Clara, CA · On-site
$142K - $192K/yr
At the center of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software. We are hiring a Senior Software Engineer to advance the ...
Senior Software Engineer, CUDA Python Core Libraries
Santa Clara, CA · On-site
$142K - $192K/yr
At the center of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software. We are hiring a Senior Software Engineer to advance the ...
At the center of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software. We are hiring a Senior Software Engineer to develop the ...
At the center of this platform are CUDA Core Libraries that enable developers to build fast, reliable, and scalable GPU-accelerated software. We are hiring a Senior Software Engineer to develop the ...
Senior Software Engineer - CUDA and Unified Memory
Santa Clara, CA · On-site
$143K - $189K/yr
CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, and ...
Senior Software Engineer - CUDA and Unified Memory
Santa Clara, CA · On-site
$143K - $189K/yr
CUDA defines a unified programming model across a range of system configurations and hardware capabilities. To accomplish this, the CUDA driver interacts with GPU hardware, kernel mode drivers, and ...
CUDA information
See salary details
$111.5K - $120.1K
0% of jobs
$120.1K - $128.7K
0% of jobs
$128.7K - $137.3K
0% of jobs
$137.3K - $145.9K
4% of jobs
$145.9K - $154.5K
0% of jobs
$154.5K - $163K
0% of jobs
$163K - $171.6K
0% of jobs
$171.6K - $180.2K
2% of jobs
$180.2K - $188.8K
0% of jobs
$188.8K - $197.4K
0% of jobs
$199.1K is the 25th percentile. Wages below this are outliers.
$197.4K - $206K
94% of jobs
$111.5K
$206K
How much do cuda jobs pay per year?
What are some common challenges faced when working as a CUDA developer, and how can they be addressed?
What are the key skills and qualifications needed to thrive as a CUDA developer, and why are they important?
What is the difference between Cuda vs GPU Developer?
| Aspect | Cuda | GPU Developer |
|---|---|---|
| Required Credentials | Knowledge of CUDA programming, often with a background in computer science or engineering | Experience with GPU programming, CUDA, OpenCL, or similar; often requires a degree in computer science or related fields |
| Work Environment | Primarily focused on developing and optimizing CUDA-based applications for NVIDIA GPUs | Designing, developing, and maintaining GPU-accelerated applications across various platforms and hardware |
| Industry Usage | Used mainly in high-performance computing, AI, and scientific research involving NVIDIA GPUs | Applied across gaming, scientific computing, AI, and multimedia industries |
In summary, CUDA is a specialized skill set focused on programming NVIDIA GPUs using CUDA, while a GPU Developer has a broader role that may include using various GPU programming tools and working across multiple platforms. CUDA is a subset of the skills a GPU Developer might possess, making them closely related but distinct roles.
What is a CUDA developer?
A CUDA job typically involves developing, optimizing, and implementing parallel computing applications using NVIDIA's CUDA platform. CUDA (Compute Unified Device Architecture) enables developers to leverage the power of GPUs for high-performance computing tasks such as deep learning, simulations, and scientific computing. Professionals in this role often work with C, C++, or Python, using CUDA libraries and frameworks to accelerate processing. Strong knowledge of parallel programming, memory management, and GPU architecture is essential for success in this field.
What cities are hiring for Cuda jobs?
Cities with the most Cuda job openings:
What are the most commonly searched types of Cuda jobs?
The most popular types of Cuda jobs are:
What states have the most Cuda jobs?
States with the most job openings for Cuda jobs include:
What job categories do people searching Cuda jobs look for?
The top searched job categories for Cuda jobs are:

$143K - $189K/yr
Full-time
Re-posted 2 days ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
7th of 244 rated software companies
Job description
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The GPU, our invention, serves as the visual cortex of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions from artificial intelligence to autonomous cars.
We are looking for a motivated Deep Learning engineer to bring advanced CUDA features and Distributed Runtime technologies into AI stacks, including PyTorch, TRT-LLM, vLLM, SGLang, JAX, etc. You will be working with the team that created core CUDA features and runtimes for scaling Deep Learning and HPC applications. Your customers will have diverse multi-GPU demands, ranging from training on scales up to 100K GPUs to inference down at microsecond latency. CUDA features improve both productivity and performance of AI applications. Your work in AI toolkits will accelerate enabling those for the community. This is an outstanding opportunity for someone with an AI background to advance the state of the art in this space. Are you ready to contribute to the development of innovative technologies and help realize NVIDIA's vision?
What you will be doing:
Integrate new CUDA features and Runtime abstractions in AI frameworks: from PoC to performance analysis to production
Perform deep analysis of AI workloads and frameworks to identify requirements and opportunities to innovate in the lower layers of the stack. Collaborate hands-on with teams working on the latest AI models.
Own and drive improvements in the AI Compiler-Runtime interface to build speed-of-light multi-GPU multi-node solutions.
Design fault-tolerant and elastic solutions for large-scale or dynamic AI workloads.
Influence the roadmap of core CUDA to facilitate building next-gen DL frameworks.
Collaborate with a very dynamic team across multiple time zones.
Collaborate closely with AI researchers, HW and SW architects, kernel and compiler authors and CUDA driver experts to co-design systems and frameworks that enhance performance and programmability.
Develop exploratory tools and runtime systems to profile and accelerate new paradigms in deep learning.
Write clean, effective, and maintainable code, ensuring exploratory prototypes can smoothly transition into open-source releases, upstream framework integrations, internal tools, or closed-source commercial products.
What we need to see:
BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field (or equivalent experience).
8+ years of relevant industry experience or equivalent academic experience after completed degree.
Development experience with Deep Learning Frameworks such PyTorch, JAX, and Inference Engines such as TRT-LLM, vLLM, SGLang
Rapid prototyping and development with Python, C++, CUDA or related DSLs
Solid grasp of AI models, parallelisms, and/or compiler technologies (e.g. torch.compile)
Experience conducting performance benchmarking on AI clusters. Familiarity with at least one performance profiler toolchain (PyTorch profiler, NVIDIA Nsight Systems)
Understanding of HPC/AI communication concepts
Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)
Adaptability and passion to learn new frameworks and tools
Flexibility to work and communicate effectively across different teams and timezones
Ways to stand out from the crowd:
Deep expertise in the performance internals and execution graphs of major deep learning autograd, training and inference frameworks (e.g., PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, Megatron, MaxText, etc.).
Hands-on experience with CUDA, specific communication libraries (e.g., NCCL, MPI, UCX) and distributed machine learning techniques (e.g., pipeline parallelism, tensor parallelism).
Expertise in one or more of these areas: Training, Distributed inference, MoE, Reinforcement Learning, kernel authoring (on CUDA, Triton, cuTe, etc).
Background in deep learning compilers, both graph-level and codegen (e.g., Triton, XLA, torch compile)
Experience with programming for compute & communication overlap in distributed runtime
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993