1

Intern Cuda Developer Jobs in Milpitas, CA (NOW HIRING)

next page

Showing results 1-20

Intern Cuda Developer information

See Milpitas, CA salary details

$13

$26

$45

How much do intern cuda developer jobs pay per hour?

As of Sep 6, 2026, the average hourly pay for intern cuda developer in Milpitas, CA is $26.68, according to ZipRecruiter salary data. Most workers in this role earn between $21.59 and $28.32 per hour, depending on experience, location, and employer.

What is the difference between Intern Cuda Developer vs Intern Software Engineer?

AspectIntern Cuda DeveloperIntern Software Engineer
Required SkillsCUDA programming, C/C++, parallel computingGeneral programming, algorithms, software development
Work EnvironmentTech companies, research labs focusing on GPU computingVarious industries including tech, finance, healthcare
CertificationsNone mandatory, but CUDA certifications helpfulNone mandatory, general software certifications beneficial

Intern Cuda Developers focus on GPU programming and parallel computing using CUDA, often in specialized tech or research settings. Intern Software Engineers have broader programming responsibilities across various platforms and industries. While both roles require programming skills, CUDA knowledge is specific to Intern Cuda Developers, making their work more specialized in high-performance computing.

What cities near Milpitas, CA are hiring for Intern Cuda Developer jobs?

Cities near Milpitas, CA with the most Intern Cuda Developer job openings:

Infographic showing various Intern Cuda Developer job openings in Milpitas, CA as of August 2026, with employment types broken down into 74% Full Time, 10% Part Time, 2% Temporary, and 14% Contract. Highlights an 80% Physical, 6% Hybrid, and 14% Remote job distribution, with an average salary of $55,496 per year, or $26.7 per hour.

Inference Optimization Intern - Performance Modeling

Institute of Foundation Models

Sunnyvale, CA โ€ข On-site

Internship

Re-posted 12 days ago


Job description

About the Institute of Foundation Models
 
The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting-edge foundation models while pushing the limits of high-performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real-world impact.
As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures. This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.
Key Responsibilities
This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.
Responsibilities include:
  • Develop analytical performance models for GPU kernels and inference workloads.
  • Build and validate a simulator to estimate theoretical hardware performance limits.
  • Compare measured kernel performance against architectural peak throughput.
  • Identify performance bottlenecks in compute, memory, communication, and scheduling.
  • Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
  • Investigate PTX and SASS code generation to understand low-level execution behavior.
  • Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
  • Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
  • Design profiling methodologies for Hopper and Blackwell architectures.
  • Document findings and provide actionable recommendations for performance improvements.
Academic Qualifications
Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
Preferred Qualifications
  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.
  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
  • Knowledge of PTX, SASS, and low-level GPU execution.
  • Experience optimizing CUDA kernels for throughput and latency.
  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Strong programming skills in C++, CUDA, and Python.
Desired Skills
  • Performance engineering mindset.
  • Strong analytical and debugging abilities.
  • Interest in AI systems, inference optimization, and hardware-software co-design.
  • Ability to work independently on research and engineering challenges.
  • Excellent written and verbal communication skills.