1

Internship Gpu Programming Jobs in California (NOW HIRING)

... internship experience and candidates with 2-5 years of industry experience in performance or systems engineering. * Hands-on experience running AI/ML workloads on GPU clusters - including ...

Showing results 21-40

Internship Gpu Programming information

What are the key skills and qualifications needed to thrive as an internship in GPU programming, and why are they important?

To thrive as an Internship GPU Programming, you need a solid background in computer science, mathematics, and programming languages such as C++ and Python, often supported by coursework or personal projects in parallel computing. Familiarity with GPU programming frameworks like CUDA or OpenCL and version control systems (e.g., Git) is typically expected. Strong analytical thinking, attention to detail, and effective communication help interns collaborate with teams and troubleshoot complex issues. These skills and qualities are essential for efficiently developing, optimizing, and debugging GPU-accelerated applications in a fast-paced, technical environment.

What types of projects or tasks can an intern expect to work on in a GPU programming internship?

As a GPU programming intern, you can expect to work on tasks such as optimizing existing code for GPU acceleration, developing parallel algorithms using CUDA or OpenCL, and assisting in the profiling and debugging of GPU applications. Interns often collaborate with researchers and software engineers to implement new features or improve the performance of computational workflows. You may also contribute to documentation and testing, gaining exposure to real-world applications in fields like machine learning, scientific computing, or graphics rendering.

What is an internship in GPU programming?

An Internship in GPU Programming is a temporary position, often held by students or recent graduates, where individuals gain hands-on experience working with Graphics Processing Units (GPUs) to develop, optimize, and accelerate software applications. Interns typically work on projects involving parallel computing, machine learning, graphics rendering, or scientific simulations using programming languages such as CUDA or OpenCL. These internships provide an opportunity to learn from experienced engineers, contribute to real-world projects, and develop specialized skills that are valuable in technology and research industries.
What are the most commonly searched types of Gpu Programming jobs in California? The most popular types of Gpu Programming jobs in California are:
What job categories do people searching Internship Gpu Programming jobs in California look for? The top searched job categories for Internship Gpu Programming jobs in California are:
What cities in California are hiring for Internship Gpu Programming jobs? Cities in California with the most Internship Gpu Programming job openings:
Infographic showing various Internship Gpu Programming job openings in California as of August 2026, with employment types broken down into 1% Internship, 79% Full Time, 13% Part Time, 2% Temporary, 4% Contract, and 1% Nights. Highlights an 89% Physical, 3% Hybrid, and 8% Remote job distribution.

Inference Optimization Intern - Performance Modeling

Institute of Foundation Models

Sunnyvale, CA • On-site

Internship

Re-posted 19 days ago


Job description

About the Institute of Foundation Models
The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting-edge foundation models while pushing the limits of high-performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real-world impact.
As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures. This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.
Key Responsibilities
This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.
Responsibilities include:
  • Develop analytical performance models for GPU kernels and inference workloads.
  • Build and validate a simulator to estimate theoretical hardware performance limits.
  • Compare measured kernel performance against architectural peak throughput.
  • Identify performance bottlenecks in compute, memory, communication, and scheduling.
  • Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
  • Investigate PTX and SASS code generation to understand low-level execution behavior.
  • Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
  • Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
  • Design profiling methodologies for Hopper and Blackwell architectures.
  • Document findings and provide actionable recommendations for performance improvements.

Academic Qualifications
Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
Preferred Qualifications
  • Experience with CUDA programming and GPU kernel development.
  • Understanding of NVIDIA GPU architecture and memory hierarchy.
  • Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
  • Knowledge of PTX, SASS, and low-level GPU execution.
  • Experience optimizing CUDA kernels for throughput and latency.
  • Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Strong programming skills in C++, CUDA, and Python.

Desired Skills
  • Performance engineering mindset.
  • Strong analytical and debugging abilities.
  • Interest in AI systems, inference optimization, and hardware-software co-design.
  • Ability to work independently on research and engineering challenges.
  • Excellent written and verbal communication skills.