1

Cuda Kernel Engineer Jobs in Raleigh, NC (NOW HIRING)

... CUDA library, and DL frameworks teams to ensure fast, functional, and timely kernel delivery to ... Strong proficiency in C++ programming and software design, including debugging, performance ...

Senior Solution Engineer, Compute Systems

Durham, NC · On-site

$53.50 - $68.75/hr

... kernel, drivers * Proficient in Python programming with the ability to build custom tools Ways to stand out from the crowd: * Background with parallel programming or GPU acceleration (e.g., CUDA)

Senior Solution Engineer, Networking

Durham, NC

$53.50 - $68.75/hr

Strong system software (firmware, BIOS, kernel, driver, operating system) expertise * Professional ... Knowledge of HPC performance test tools and NVIDIA AI stacks (NCCL, MPI, DOCA, CUDA) Widely ...

Support CPU architects and performance engineers in their use of functional models, performance ... Experience with Linux kernel bringup and debug * Familiarity with CUDA * Experience with CPU/GPU ...

Cuda Kernel Engineer information

What is a CUDA Kernel Engineer?

Cuda Kernel Engineers are specialized software developers who design, implement, and optimize parallel computing algorithms using NVIDIA's CUDA platform. They write 'kernels,' which are functions that run on Graphics Processing Units (GPUs) to accelerate computational tasks in areas such as machine learning, scientific simulations, and graphics rendering. These engineers need strong skills in C/C++ programming, GPU architecture, and performance optimization techniques. Their work is crucial for applications that require high-speed data processing and efficient resource utilization.

What skills and qualifications are needed to be a CUDA Kernel Engineer?

To thrive as a CUDA Kernel Engineer, you need strong proficiency in C/C++ programming, parallel computing concepts, and a solid foundation in GPU architectures, typically supported by a degree in computer science or a related field. Expertise in NVIDIA CUDA toolkits, GPU profiling tools like Nsight, and familiarity with version control systems are essential. Analytical thinking, problem-solving abilities, and effective collaboration skills help engineers optimize code and work well within development teams. These skills and qualities are crucial for delivering high-performance, scalable GPU solutions in computationally intensive applications.

What are common challenges faced by CUDA Kernel Engineers when optimizing GPU code for performance?

Cuda Kernel Engineers often encounter challenges such as managing memory hierarchy efficiently, minimizing data transfer between host and device, and avoiding thread divergence. Ensuring optimal occupancy and maximizing parallelism while preventing bottlenecks like bank conflicts or uncoalesced memory access are also key concerns. Collaborating closely with software architects and data scientists is common, as solutions frequently require balancing algorithmic accuracy with hardware limitations. Addressing these challenges requires continuous profiling, testing, and iterative optimization.

What are popular job titles related to Cuda Kernel Engineer jobs in Raleigh, NC?

For Cuda Kernel Engineer jobs in Raleigh, NC, the most frequently searched job titles are:

What job categories do people searching Cuda Kernel Engineer jobs in Raleigh, NC look for?

The top searched job categories for Cuda Kernel Engineer jobs in Raleigh, NC are:

What cities near Raleigh, NC are hiring for Cuda Kernel Engineer jobs?

Cities near Raleigh, NC with the most Cuda Kernel Engineer job openings:

Senior Software Engineer, CUTLASS Kernels

Durham, NC


Nvidia
Computer and Electronic Product Manufacturing • 10K+ employees

9.6

Company rating: 9.6 out of 10

Based on 18 frontline employees who took The Breakroom Quiz

7th of 246 rated software companies

Great coworkers

People enjoy working here

Good employer


$118K - $156K/yr

Full-time

Re-posted 25 days ago


Job description

NVIDIA's high-performance computing platforms are powering the AI revolution across many applications and industries. Within our software stack, CUTLASS stands out as a popular open-source ecosystem dedicated to high-performance linear algebra and Tensor Core primitives. Since 2017, it has provided the community with C++ and Python abstractions to implement custom matrix multiply (GEMM) and related math and deep learning computations on NVIDIA GPUs.

If you are passionate about developing and optimizing math kernels to extract the highest performance out of the hardware architecture, apply to join the CUTLASS team today!

What you'll get to do:

  • Write Tensor Core-based deep learning kernels such as grouped-GEMM, attention, and convolution using CUTLASS CUDA C++ and Python DSL for Blackwell, Rubin, and future architectures.

  • Optimize kernels for peak throughput on both silicon and software performance simulators.

  • Collaborate with teams across NVIDIA including the GPU architecture, NVVM/PTX compiler, CUDA library, and DL frameworks teams to ensure fast, functional, and timely kernel delivery to customers.

What we need to see:

  • Masters or PhD degree in Computer Science, Computer Engineering, or related field (or equivalent experience).

  • 3+ years of relevant industry experience.

  • Strong proficiency in C++ programming and software design, including debugging, performance evaluation, and testing.

  • Experience with CUDA, OpenCL, HIP, SYCL, Mojo, Pallas, Triton, Mosaic, Halide, or any general-purpose or domain-specific programming language targeting highly parallel accelerators.

  • Deep understanding of computer architecture and some experience working at the assembly level.

Ways to stand out from the crowd:

  • Experience writing code specifically targeting NVIDIA Tensor Cores, particularly through PTX or CUDA/cuTile.

  • Open-source contributions to math kernel libraries or frameworks.

NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hard working people in the world working for us. If you're creative, autonomous, and love a challenge, consider joining our Deep Learning Library team and help us build the real-time, cost-effective computing platform driving our success in this exciting and quickly growing field.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 152,000 USD - 241,500 USD for Level 3, and 184,000 USD - 287,500 USD for Level 4.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 5, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US


What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom