As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures.
As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures.
Sr Software Engineer (SDE-III), Core Search
Palo Alto, CA · On-site
$144K - $189K/yr
Strong candidates will have knowledge of LLM fundamentals and some experience in GPU programming ... internship professional software development experience - 5+ years of leading design or ...
Sr Software Engineer (SDE-III), Core Search
Palo Alto, CA · On-site
$144K - $189K/yr
Strong candidates will have knowledge of LLM fundamentals and some experience in GPU programming ... internship professional software development experience - 5+ years of leading design or ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Familiarity with real-time rendering and GPU programming (CUDA, WebGL, graphics pipelines ... Internships at Netflix At Netflix, we offer a personalized experience for interns, and our aim is ...
Senior Compiler Engineer, AI Inference Platforms
Santa Clara, CA · On-site
$122K - $168K/yr
Preferred : • Proficient in CPU and/or GPU architecture. • CUDA or OpenCL programming ... interns is a bonus. • Track record on new hardware bring-up is a plus. Company : NVIDIA is a ...
Senior Compiler Engineer, AI Inference Platforms
Santa Clara, CA · On-site
$122K - $168K/yr
Preferred : • Proficient in CPU and/or GPU architecture. • CUDA or OpenCL programming ... interns is a bonus. • Track record on new hardware bring-up is a plus. Company : NVIDIA is a ...
... in GPU architecture, proficiency in CUDA programming, programming large-scale clusters, and ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
... in GPU architecture, proficiency in CUDA programming, programming large-scale clusters, and ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
... in GPU architecture, proficiency in CUDA programming, programming large-scale clusters, and ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
... in GPU architecture, proficiency in CUDA programming, programming large-scale clusters, and ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
... in GPU architecture, proficiency in CUDA programming, programming large-scale clusters, and ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
... in GPU architecture, proficiency in CUDA programming, programming large-scale clusters, and ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
Senior Math Libraries Engineer - Sparse Linear Algebra
Santa Clara, CA · On-site
$55K - $72K/yr
... future GPU architectures • Working with library engineers, QA engineers, and interns on topics ranging from sparse BLAS operations to advanced direct and iterative sparse solvers • Working ...
Senior Math Libraries Engineer - Sparse Linear Algebra
Santa Clara, CA · On-site
$55K - $72K/yr
... future GPU architectures • Working with library engineers, QA engineers, and interns on topics ranging from sparse BLAS operations to advanced direct and iterative sparse solvers • Working ...
Software Engineer Project Intern (Model Infrastructure) - 2026 Start (BS/MS)
San Jose, CA · On-site
$45/hr
... programming skills in C++ and Python. - Solid understanding of Computer Architecture and the GPU ... Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
Software Engineer Project Intern (Model Infrastructure) - 2026 Start (BS/MS)
San Jose, CA · On-site
$45/hr
... programming skills in C++ and Python. - Solid understanding of Computer Architecture and the GPU ... Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
Designing, implementing, and optimizing direct sparse solvers for existing and future GPU architectures * Working with library engineers, QA engineers, and interns on all library development aspects ...
Designing, implementing, and optimizing direct sparse solvers for existing and future GPU architectures * Working with library engineers, QA engineers, and interns on all library development aspects ...
Senior Math Libraries Engineer - Direct Sparse Solvers
Santa Clara, CA · On-site
$143K - $189K/yr
Designing, implementing, and optimizing direct sparse solvers for existing and future GPU architectures * Working with library engineers, QA engineers, and interns on all library development aspects ...
Senior Math Libraries Engineer - Direct Sparse Solvers
Santa Clara, CA · On-site
$143K - $189K/yr
Designing, implementing, and optimizing direct sparse solvers for existing and future GPU architectures * Working with library engineers, QA engineers, and interns on all library development aspects ...
Senior Compiler Engineer, AI Inference Platforms
$143K - $189K/yr
Proficient in CPU and/or GPU architecture. CUDA or OpenCL programming experience. * Understanding ... A track record of success in mentoring early-career engineers and interns is a bonus. * Track ...
Senior Compiler Engineer, AI Inference Platforms
$143K - $189K/yr
Proficient in CPU and/or GPU architecture. CUDA or OpenCL programming experience. * Understanding ... A track record of success in mentoring early-career engineers and interns is a bonus. * Track ...
Senior Compiler Engineer, AI Inference Platforms
Santa Clara, CA · On-site
$143K - $189K/yr
Proficient in CPU and/or GPU architecture. CUDA or OpenCL programming experience. * Understanding ... A track record of success in mentoring early-career engineers and interns is a bonus. * Track ...
Senior Compiler Engineer, AI Inference Platforms
Santa Clara, CA · On-site
$143K - $189K/yr
Proficient in CPU and/or GPU architecture. CUDA or OpenCL programming experience. * Understanding ... A track record of success in mentoring early-career engineers and interns is a bonus. * Track ...
Software Engineering Intern, Dynamo - Fall 2026
Santa Clara, CA · On-site
$22.50 - $29.75/hr
An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can ... Excellent Golang, Rust and/or Python programming and software design skills, including debugging ...
Software Engineering Intern, Dynamo - Fall 2026
Santa Clara, CA · On-site
$22.50 - $29.75/hr
An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can ... Excellent Golang, Rust and/or Python programming and software design skills, including debugging ...
Software Engineering Intern, Dynamo - Fall 2026
$22.50 - $29.75/hr
An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can ... Excellent Golang, Rust and/or Python programming and software design skills, including debugging ...
Software Engineering Intern, Dynamo - Fall 2026
$22.50 - $29.75/hr
An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can ... Excellent Golang, Rust and/or Python programming and software design skills, including debugging ...
... internship experience and candidates with 2-5 years of industry experience in performance or systems engineering. * Hands-on experience running AI/ML workloads on GPU clusters - including ...
... internship experience and candidates with 2-5 years of industry experience in performance or systems engineering. * Hands-on experience running AI/ML workloads on GPU clusters - including ...
Internship Gpu Programming information
What are the key skills and qualifications needed to thrive as an internship in GPU programming, and why are they important?
What types of projects or tasks can an intern expect to work on in a GPU programming internship?
What is an internship in GPU programming?

Inference Optimization Intern - Performance Modeling
Sunnyvale, CA • On-site
Internship
Re-posted 19 days ago
Job description
The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting-edge foundation models while pushing the limits of high-performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real-world impact.
As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures. This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.
Key Responsibilities
This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.
Responsibilities include:
- Develop analytical performance models for GPU kernels and inference workloads.
- Build and validate a simulator to estimate theoretical hardware performance limits.
- Compare measured kernel performance against architectural peak throughput.
- Identify performance bottlenecks in compute, memory, communication, and scheduling.
- Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
- Investigate PTX and SASS code generation to understand low-level execution behavior.
- Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
- Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
- Design profiling methodologies for Hopper and Blackwell architectures.
- Document findings and provide actionable recommendations for performance improvements.
Academic Qualifications
Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
Preferred Qualifications
- Experience with CUDA programming and GPU kernel development.
- Understanding of NVIDIA GPU architecture and memory hierarchy.
- Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
- Knowledge of PTX, SASS, and low-level GPU execution.
- Experience optimizing CUDA kernels for throughput and latency.
- Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
- Experience with deep learning frameworks such as PyTorch or TensorFlow.
- Strong programming skills in C++, CUDA, and Python.
Desired Skills
- Performance engineering mindset.
- Strong analytical and debugging abilities.
- Interest in AI systems, inference optimization, and hardware-software co-design.
- Ability to work independently on research and engineering challenges.
- Excellent written and verbal communication skills.