1

Cuda Kernel Engineer Jobs in Arizona (NOW HIRING)

Sr. Machine Learning Engineer

Phoenix, AZ · On-site

$130K - $150K/yr

Sr. Machine Learning Engineer Salary Range: $130k to $150k Our client is seeking a Sr. Machine ... Responsibilities include eliminating hardware bottlenecks through CUDA kernel tuning and GPU ...

Strong expertise in CUDA programming and GPU kernel optimization . * Advanced C++ development skills. * Hands-on experience with GLSL and WebGPU . * Experience using GPU profiling tools such as ...

CUDA Developer - Remote

Phoenix, AZ · Remote

$60 - $100/hr

Strong expertise in CUDA programming and GPU kernel optimization . * Advanced C++ development skills. * Hands-on experience with GLSL and WebGPU . * Experience using GPU profiling tools such as ...

Strong expertise in CUDA programming and GPU kernel optimization . * Advanced C++ development skills. * Hands-on experience with GLSL and WebGPU . * Experience using GPU profiling tools such as ...

Cuda Kernel Engineer information

What is a CUDA Kernel Engineer?

Cuda Kernel Engineers are specialized software developers who design, implement, and optimize parallel computing algorithms using NVIDIA's CUDA platform. They write 'kernels,' which are functions that run on Graphics Processing Units (GPUs) to accelerate computational tasks in areas such as machine learning, scientific simulations, and graphics rendering. These engineers need strong skills in C/C++ programming, GPU architecture, and performance optimization techniques. Their work is crucial for applications that require high-speed data processing and efficient resource utilization.

What skills and qualifications are needed to be a CUDA Kernel Engineer?

To thrive as a CUDA Kernel Engineer, you need strong proficiency in C/C++ programming, parallel computing concepts, and a solid foundation in GPU architectures, typically supported by a degree in computer science or a related field. Expertise in NVIDIA CUDA toolkits, GPU profiling tools like Nsight, and familiarity with version control systems are essential. Analytical thinking, problem-solving abilities, and effective collaboration skills help engineers optimize code and work well within development teams. These skills and qualities are crucial for delivering high-performance, scalable GPU solutions in computationally intensive applications.

What are common challenges faced by CUDA Kernel Engineers when optimizing GPU code for performance?

Cuda Kernel Engineers often encounter challenges such as managing memory hierarchy efficiently, minimizing data transfer between host and device, and avoiding thread divergence. Ensuring optimal occupancy and maximizing parallelism while preventing bottlenecks like bank conflicts or uncoalesced memory access are also key concerns. Collaborating closely with software architects and data scientists is common, as solutions frequently require balancing algorithmic accuracy with hardware limitations. Addressing these challenges requires continuous profiling, testing, and iterative optimization.

What are popular job titles related to Cuda Kernel Engineer jobs in Arizona?

For Cuda Kernel Engineer jobs in Arizona, the most frequently searched job titles are:

What job categories do people searching Cuda Kernel Engineer jobs in Arizona look for?

The top searched job categories for Cuda Kernel Engineer jobs in Arizona are:

What cities in Arizona are hiring for Cuda Kernel Engineer jobs?

Cities in Arizona with the most Cuda Kernel Engineer job openings:

Sr. Machine Learning Engineer

Prosum Inc.

Phoenix, AZ • On-site

$130K - $150K/yr

Other

Posted 18 days ago


Job description

Job Description
Sr. Machine Learning Engineer
Salary Range: $130k to $150k
Our client is seeking a Sr. Machine Learning Engineering for a direct hire role to sit in North Phoenix, AZ or Hillsboro, OR. This role will be onsite 4 days a week and 1 remote day.
JOB SUMMARY
The role of Senior Machine Learning Engineer will architect and optimize real-time, high-throughput, and ultra-low latency image pipelines for next-generation Mask Inspection Tools. Responsibilities include eliminating hardware bottlenecks through CUDA kernel tuning and GPU parallel computing, ensuring deep learning models and CV algorithms seamlessly processing massive, high-bandwidth streaming data at production scale.
ESSENTIAL DUTIES AND RESPONSIBILITIES
High-Performance Computing Pipeline Architecture
  • Design, implement, and optimize high-throughput, low-latency image processing pipelines for real-time optical inspection and machine vision systems.
  • Develop scalable architectures capable of processing large volumes of imaging data while meeting stringent latency and reliability requirements.
  • Profile and optimize system performance across CPU, GPU, memory, and I/O subsystems.
GPU Acceleration
  • Design, develop, and optimize CUDA kernels to accelerate deep learning inference and classical computer vision algorithms.
  • Maximize GPU utilization through efficient memory management, kernel optimization, and parallel programming techniques.
  • Evaluate and implement performance improvements using NVIDIA GPU technologies and profiling tools.
Model Deployment & Optimization
  • Optimize, quantize, and deploy machine learning models using TensorRT, ONNX Runtime, or similar inference frameworks.
  • Integrate AI models into production-grade C++ and Python applications.
  • Improve inference throughput, latency, and resource utilization while maintaining model accuracy.
  • Develop automated deployment and validation pipelines for machine learning models.
Concurrency & Systems Optimization
  • Architect and implement multi-threaded, high-concurrency software components for data acquisition, buffering, streaming, and real-time processing.
  • Design robust synchronization and communication mechanisms between hardware interfaces and AI processing pipelines.
  • Optimize end-to-end system performance for deterministic, real-time execution.
Cross-Functional Collaboration
  • Partner with machine learning scientists, computer vision engineers, hardware engineers, and software developers to deliver integrated AI solutions.

Please view our Privacy Policy.