1

Cuda Kernel Engineer Jobs in Missouri (NOW HIRING)

Cuda Kernel Engineer information

What is a CUDA Kernel Engineer?

Cuda Kernel Engineers are specialized software developers who design, implement, and optimize parallel computing algorithms using NVIDIA's CUDA platform. They write 'kernels,' which are functions that run on Graphics Processing Units (GPUs) to accelerate computational tasks in areas such as machine learning, scientific simulations, and graphics rendering. These engineers need strong skills in C/C++ programming, GPU architecture, and performance optimization techniques. Their work is crucial for applications that require high-speed data processing and efficient resource utilization.

What skills and qualifications are needed to be a CUDA Kernel Engineer?

To thrive as a CUDA Kernel Engineer, you need strong proficiency in C/C++ programming, parallel computing concepts, and a solid foundation in GPU architectures, typically supported by a degree in computer science or a related field. Expertise in NVIDIA CUDA toolkits, GPU profiling tools like Nsight, and familiarity with version control systems are essential. Analytical thinking, problem-solving abilities, and effective collaboration skills help engineers optimize code and work well within development teams. These skills and qualities are crucial for delivering high-performance, scalable GPU solutions in computationally intensive applications.

What are common challenges faced by CUDA Kernel Engineers when optimizing GPU code for performance?

Cuda Kernel Engineers often encounter challenges such as managing memory hierarchy efficiently, minimizing data transfer between host and device, and avoiding thread divergence. Ensuring optimal occupancy and maximizing parallelism while preventing bottlenecks like bank conflicts or uncoalesced memory access are also key concerns. Collaborating closely with software architects and data scientists is common, as solutions frequently require balancing algorithmic accuracy with hardware limitations. Addressing these challenges requires continuous profiling, testing, and iterative optimization.

What are popular job titles related to Cuda Kernel Engineer jobs in Missouri?

For Cuda Kernel Engineer jobs in Missouri, the most frequently searched job titles are:

What job categories do people searching Cuda Kernel Engineer jobs in Missouri look for?

The top searched job categories for Cuda Kernel Engineer jobs in Missouri are:

What cities in Missouri are hiring for Cuda Kernel Engineer jobs?

Cities in Missouri with the most Cuda Kernel Engineer job openings:

Infographic showing various Cuda Kernel Engineer job openings in Missouri as of August 2026, with employment types broken down into 100% Full Time. Highlights an 60% In-person, and 40% Remote job distribution.

Software Engineer, CUDA Deep Learning Systems

Jobtailor

California, MO • On-site

$140 - $210/hr

Other

Posted 10 days ago


Job description

  • Explore, research, and prototype systems optimizations for advanced deep learning models at the intersection of high-level deep learning frameworks and low-level CUDA.
  • Architect and optimize distributed computing systems from single-node to cluster-scale supercomputing environments.
  • Design, implement, and optimize custom high-performance CUDA kernels for emerging neural network architectures and workloads.
  • Analyze hardware-software interactions to identify and resolve performance bottlenecks in training and inference pipelines.
  • Collaborate with AI researchers, hardware and software architects, kernel and compiler authors, and CUDA driver experts to co-design systems and algorithms.
  • Develop exploratory tools and runtime systems to profile and accelerate new deep learning paradigms.
  • Write clean, effective, and maintainable code and transition prototypes into open-source releases, framework integrations, internal tools, or commercial products.
Requirements
  • BS, MS, or PhD degree in Computer Science, Computer Engineering, Electrical Engineering, or related field, or equivalent experience.
  • 2+ years of relevant industry experience or equivalent academic experience after degree achievement.
  • Strong proficiency in C++ and Python programming.
  • Solid background in deep learning fundamentals, focused on transformers.
  • Strong understanding of distributed computing, multi-node scaling, and cluster-scale performance challenges.
  • Proven experience in systems programming, computer architecture, and low-level systems performance optimization.
  • Familiarity with GPU deep learning accelerator architectures.
  • Hands-on experience with CUDA programming, kernel optimization, and workload profiling.
  • Experience profiling and optimizing generative AI models, including large language models.
  • Research background in machine learning systems or adjacent fields.
  • Experience profiling and optimizing innovative vision models, generative AI architectures, or diffusion models.
  • Track record of initiative and willingness to deep-dive on problems across the stack.
  • Preferred experience with PyTorch, JAX, TensorRT, vLLM, sgLang, Nemo, or Megatron internals and execution graphs.
  • Preferred hands-on experience with NCCL, MPI, or UCX and distributed machine learning techniques.
  • Preferred knowledge of numerical methods and low-precision arithmetic such as NVFP4, MXFP4, FP8, and INT8.
  • Preferred background in deep learning compilers and ML systems, including Triton, XLA, and torch.compile.
  • Preferred experience with highly parallel or reinforcement-learning-style simulation environments.
  • Preferred experience designing and implementing agentic AI systems for complex systems and infrastructure problems.
Core Competencies

Demonstrates expertise in CUDA programming, high-performance computing, and deep learning optimization, with a strong foundation in distributed systems and machine learning architectures. Capable of collaborating with cross-functional teams to design and implement innovative solutions for advanced AI models.

Highest-signal resume keywords
  • CUDA Programming
  • C++ and Python Proficiency
  • Deep Learning Optimization
  • Distributed Computing Systems
  • Performance Bottleneck Analysis
ATS Optimization Keywords Hard Skills
  • CUDA
  • C++
  • Python
  • Deep Learning Fundamentals
  • Systems Programming
  • Computer Architecture
  • Kernel Optimization
  • Workload Profiling
  • Generative AI Models
  • Transformers
Soft Skills
  • Collaboration
  • Initiative
  • Problem-Solving
Industry Keywords
  • Machine Learning Systems
  • High-Performance Computing
  • AI Research
  • Neural Network Architectures
  • Cluster-Scale Performance
Tools & Technologies
  • PyTorch
  • JAX
  • TensorRT
  • NCCL
  • MPI
  • UCX
  • Triton
  • XLA
  • Torch.compile
  • SgLang
#J-18808-Ljbffr