1

Cuda Kernel Engineer Jobs (NOW HIRING)

Review GPU kernel implementations and identify bottlenecks without requiring extensive background in the underlying algorithms * Write, modify, and reason about C++17, Python, and GPU programming ...

System Software Engineer - CUDA Chips

Santa Clara, CA · On-site

$203K - $240K/yr

... kernel mode drivers, and the operating system. This role incorporates strong system software ... CUDA APIs and programming model, while driving development efforts across multiple teams Write ...

We're looking for a CUDA Engineer to write and optimise the low-level GPU code that powers our ... Apply kernel fusion to reduce memory round-trips and launch overhead across inference pipelines

$90K - $119K/yr

We are hiring senior engineers to work on the CUDA driver, a core component of our platform for ... Background with kernel mode development * Experience with Windows, Linux, or macOS driver ...

Showing results 21-40

Cuda Kernel Engineer information

What is a CUDA Kernel Engineer?

Cuda Kernel Engineers are specialized software developers who design, implement, and optimize parallel computing algorithms using NVIDIA's CUDA platform. They write 'kernels,' which are functions that run on Graphics Processing Units (GPUs) to accelerate computational tasks in areas such as machine learning, scientific simulations, and graphics rendering. These engineers need strong skills in C/C++ programming, GPU architecture, and performance optimization techniques. Their work is crucial for applications that require high-speed data processing and efficient resource utilization.

What skills and qualifications are needed to be a CUDA Kernel Engineer?

To thrive as a CUDA Kernel Engineer, you need strong proficiency in C/C++ programming, parallel computing concepts, and a solid foundation in GPU architectures, typically supported by a degree in computer science or a related field. Expertise in NVIDIA CUDA toolkits, GPU profiling tools like Nsight, and familiarity with version control systems are essential. Analytical thinking, problem-solving abilities, and effective collaboration skills help engineers optimize code and work well within development teams. These skills and qualities are crucial for delivering high-performance, scalable GPU solutions in computationally intensive applications.

What are common challenges faced by CUDA Kernel Engineers when optimizing GPU code for performance?

Cuda Kernel Engineers often encounter challenges such as managing memory hierarchy efficiently, minimizing data transfer between host and device, and avoiding thread divergence. Ensuring optimal occupancy and maximizing parallelism while preventing bottlenecks like bank conflicts or uncoalesced memory access are also key concerns. Collaborating closely with software architects and data scientists is common, as solutions frequently require balancing algorithmic accuracy with hardware limitations. Addressing these challenges requires continuous profiling, testing, and iterative optimization.
More about Cuda Kernel Engineer jobs

What cities are hiring for Cuda Kernel Engineer jobs?

Cities with the most Cuda Kernel Engineer job openings:

What states have the most Cuda Kernel Engineer jobs?

States with the most job openings for Cuda Kernel Engineer jobs include:

What are popular job titles related to Cuda Kernel Engineer jobs?

For Cuda Kernel Engineer jobs, the most frequently searched job titles are:

Infographic showing various Cuda Kernel Engineer job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 92% Full Time, 3% Part Time, and 4% Contract. Highlights an 85% Physical, 4% Hybrid, and 11% Remote job distribution.

Kernel Engineer - New Grad

Sunnyvale, CA • On-site

Cerebras Systems
Computer and Peripheral Equipment Manufacturing • 201 - 500 employees

Full-time

Posted 14 days ago


Key responsibilities

  • Help design and implement machine learning and linear algebra kernels for the Cerebras Wafer-Scale Engine.

  • Develop and debug high-performance kernel routines using low-level programming techniques and the Cerebras Software Language.

  • Apply parallel programming algorithms to map computational workloads efficiently onto the Cerebras architecture.


Job description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services.
This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
About the Role
As a Kernel Engineer at Cerebras, you will develop high-performance software at the intersection of hardware and software for cutting-edge artificial intelligence and high-performance computing workloads.
You will help implement, optimize, and validate machine learning and linear algebra operations for the Cerebras Wafer-Scale Engine, our custom massively parallel processor architecture. Working alongside experienced kernel, compiler, performance, and hardware engineers, you will learn how algorithms are mapped to specialized hardware and contribute to software that maximizes compute utilization and system performance.
You will be part of a team responsible for the design, development, performance tuning, and validation of foundational ML and HPC kernels. This is an excellent opportunity for a new graduate who is interested in computer architecture, parallel programming, low-level software, and machine learning systems.
Responsibilities
  • Help design and implement machine learning and linear algebra kernels for the Cerebras Wafer-Scale Engine.
  • Develop and debug high-performance kernel routines using low-level programming techniques and the Cerebras Software Language, a custom C-like language.
  • Apply parallel programming algorithms to map computational workloads efficiently onto the Cerebras architecture.
  • Use mathematical analysis, performance data, and profiling tools to evaluate kernel behavior and inform design decisions.
  • Identify and investigate correctness, performance, and hardware utilization issues.
  • Develop unit tests and system-level validation methodologies to verify the functionality and performance of kernel libraries.
  • Collaborate with kernel, compiler, performance, and hardware engineers to improve software and system performance.
  • Study emerging machine learning workloads and contribute to the evolution of the kernel library.
  • Participate in code reviews, technical discussions, and software development processes.
  • Build an understanding of the Cerebras architecture, instruction set, memory system, and communication model.
Minimum Skills & Qualifications
  • Bachelor's, Master's, or PhD in Computer Science, Computer Engineering, Electrical Engineering, Mathematics, or a related field.
  • Strong programming fundamentals in C++ and familiarity with Python.
  • Understanding of foundational computer architecture concepts such as processors, memory hierarchies, instruction execution, or data movement.
  • Knowledge of data structures, algorithms, and software development fundamentals.
  • Experience debugging software through coursework, internships, research, co-op placements, or technical projects.
  • Strong analytical and problem-solving skills.
  • Interest in low-level software, parallel computing, performance optimization, or hardware/software co-design.
  • Ability to learn unfamiliar systems and collaborate effectively within a technical team.
Preferred Skills & Qualifications
  • Research, internships, or projects involving kernel development, compilers, computer architecture, HPC, or systems programming.
  • Familiarity with parallel algorithms, multithreaded programming, or distributed memory systems.
  • Exposure to programming accelerators such as GPUs, FPGAs, or other specialized processors.
  • Experience with low-level programming, assembly language, CUDA, OpenCL, or a domain-specific language.
  • Familiarity with machine learning concepts, neural networks, or frameworks such as PyTorch or TensorFlow.
  • Exposure to numerical computing, linear algebra, or HPC kernels.
  • Experience using profiling, benchmarking, or performance analysis tools.
  • Familiarity with library or API development practices.

Why Join Cerebras
People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we've reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:
  1. Build a breakthrough AI platform beyond the constraints of the GPU.
  2. Publish and open source their cutting-edge AI research.
  3. Work on one of the fastest AI supercomputers in the world.
  4. Enjoy job stability with startup vitality.
  5. Our simple, non-corporate work culture that respects individual beliefs.

Find out more about what it's like to work at Cerebras here!
Apply today and become part of the forefront of groundbreaking advancements in AI!
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.