As part of the team, interns work alongside world-class researchers and performance engineers to ... Experience with CUDA programming and GPU kernel development. * Understanding of NVIDIA GPU ...
As part of the team, interns work alongside world-class researchers and performance engineers to ... Experience with CUDA programming and GPU kernel development. * Understanding of NVIDIA GPU ...
NVIDIA has pioneered programmable GPUs and the CUDA language, and is a world leader in high ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
NVIDIA has pioneered programmable GPUs and the CUDA language, and is a world leader in high ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
NVIDIA has pioneered programmable GPUs and the CUDA language, and is a world leader in high ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
NVIDIA has pioneered programmable GPUs and the CUDA language, and is a world leader in high ... Our internship hourly rates are a standard pay based on the position, your location, year in school ...
... Interns in San Francisco. You'll work on fundamental problems in LLM-based agentic systems and ... Proficiency in CUDA programming and custom kernel development for LLM operations * Background in ...
... Interns in San Francisco. You'll work on fundamental problems in LLM-based agentic systems and ... Proficiency in CUDA programming and custom kernel development for LLM operations * Background in ...
GPU/AI Application System Software Engineer Intern (System Technologies and Engineering) - 2026 Summ
San Jose, CA · On-site
$45/hr
Internships at TikTok aim to offer students industry exposure and hands-on experience. Turn your ... CUDA programming - Linux kernel development experience, such as networking and device drivers etc ...
GPU/AI Application System Software Engineer Intern (System Technologies and Engineering) - 2026 Summ
San Jose, CA · On-site
$45/hr
Internships at TikTok aim to offer students industry exposure and hands-on experience. Turn your ... CUDA programming - Linux kernel development experience, such as networking and device drivers etc ...
Excellent parallel C++ programming, with familiarity in CUDA programming * Deep understanding of ... The hourly rate for our interns is 30 USD - 94 USD. You will also be eligible for Intern benefits.
Excellent parallel C++ programming, with familiarity in CUDA programming * Deep understanding of ... The hourly rate for our interns is 30 USD - 94 USD. You will also be eligible for Intern benefits.
Excellent parallel C++ programming, with familiarity in CUDA programming * Deep understanding of ... The hourly rate for our interns is 30 USD - 94 USD. You will also be eligible for Internbenefits.
Excellent parallel C++ programming, with familiarity in CUDA programming * Deep understanding of ... The hourly rate for our interns is 30 USD - 94 USD. You will also be eligible for Internbenefits.
Research Intern, Inference (Fall 2026)
San Francisco, CA · On-site
$58 - $63/hr
... CUDA programming (for kernel development) * Understanding of model optimization techniques and hardware acceleration approaches * Contributions to open-source machine learning projects Internship ...
Research Intern, Inference (Fall 2026)
San Francisco, CA · On-site
$58 - $63/hr
... CUDA programming (for kernel development) * Understanding of model optimization techniques and hardware acceleration approaches * Contributions to open-source machine learning projects Internship ...
... CUDA programming (for kernel development) * Understanding of model optimization techniques and hardware acceleration approaches * Contributions to open-source machine learning projects Internship ...
... CUDA programming (for kernel development) * Understanding of model optimization techniques and hardware acceleration approaches * Contributions to open-source machine learning projects Internship ...
... CUDA kernels and TensorRT optimization techniques - Partner with scientists to analyze model ... non-internship professional software development experience, or Bachelor's degree in computer ...
... CUDA kernels and TensorRT optimization techniques - Partner with scientists to analyze model ... non-internship professional software development experience, or Bachelor's degree in computer ...
NVIDIA is looking for Java engineering interns to work on cuVS, a suite of open source software ... This is a great chance to take advantage of your Java CUDA/C++ skills and make a huge impact across ...
NVIDIA is looking for Java engineering interns to work on cuVS, a suite of open source software ... This is a great chance to take advantage of your Java CUDA/C++ skills and make a huge impact across ...
Member of Technical Staff - Simulation (Physics & Reinforcement Learning), Frontier AI Robotics
San Francisco, CA · On-site
$150K/yr
BASIC QUALIFICATIONS - 5+ years of non-internship professional software development experience - 5+ ... CUDA programming - Experience with TensorRT or similar ML optimization frameworks - Ability to ...
Member of Technical Staff - Simulation (Physics & Reinforcement Learning), Frontier AI Robotics
San Francisco, CA · On-site
$150K/yr
BASIC QUALIFICATIONS - 5+ years of non-internship professional software development experience - 5+ ... CUDA programming - Experience with TensorRT or similar ML optimization frameworks - Ability to ...
Research Scientist Intern (TikTok - NextGen Recommendation) - 2026 Start (PhD)
San Jose, CA · On-site
$60/hr
Proficiency in CUDA programming (experience with Triton) is highly desirable. - Prior research ... Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
Research Scientist Intern (TikTok - NextGen Recommendation) - 2026 Start (PhD)
San Jose, CA · On-site
$60/hr
Proficiency in CUDA programming (experience with Triton) is highly desirable. - Prior research ... Interns have day one access to health insurance, life insurance, wellbeing benefits and more.
Research Scientist, Gen AI & User Representation Learning
Palo Alto, CA · On-site +1
$100K - $130K/yr
GPU Computing * NVIDIA GPU architecture and CUDA programming fundamentals * Multi-GPU and ... internships)
Research Scientist, Gen AI & User Representation Learning
Palo Alto, CA · On-site +1
$100K - $130K/yr
GPU Computing * NVIDIA GPU architecture and CUDA programming fundamentals * Multi-GPU and ... internships)
Systems Research Engineer Intern - GPU Programming (Fall 2026)
San Francisco, CA · On-site
$58 - $63/hr
Strong background in GPU programming and parallel computing, such as CUDA and/or Triton ... Our internship dates are September 14th to December 18th. About Together AI Together AI is a ...
Systems Research Engineer Intern - GPU Programming (Fall 2026)
San Francisco, CA · On-site
$58 - $63/hr
Strong background in GPU programming and parallel computing, such as CUDA and/or Triton ... Our internship dates are September 14th to December 18th. About Together AI Together AI is a ...
Strong background in GPU programming and parallel computing, such as CUDA and/or Triton ... Our internship dates are September 14th to December 18th. About Together AI Together AI is a ...
Strong background in GPU programming and parallel computing, such as CUDA and/or Triton ... Our internship dates are September 14th to December 18th. About Together AI Together AI is a ...
Many interns receive return or full-time offers. This will be an internship for one of two ... Experience with sensor drivers, CUDA, 3D geometry, or multi-sensor fusion.
Many interns receive return or full-time offers. This will be an internship for one of two ... Experience with sensor drivers, CUDA, 3D geometry, or multi-sensor fusion.
Intern, AI Engineering
$19.75 - $25.50/hr
... Interns in San Francisco. You'll work on fundamental problems in LLM-based agentic systems and ... Proficiency in CUDA programming and custom kernel development for LLM operations * Background in ...
Quick apply
Intern, AI Engineering
$19.75 - $25.50/hr
... Interns in San Francisco. You'll work on fundamental problems in LLM-based agentic systems and ... Proficiency in CUDA programming and custom kernel development for LLM operations * Background in ...
Senior Research Scientist, Efficient Deep Learning
Santa Clara, CA · On-site
$115K - $147K/yr
Mentor interns. • Work with product groups to transfer technology. Qualifications : Required ... C++ and parallel programming (e.g., CUDA). • Hands-on experience with large-scale model training ...
Senior Research Scientist, Efficient Deep Learning
Santa Clara, CA · On-site
$115K - $147K/yr
Mentor interns. • Work with product groups to transfer technology. Qualifications : Required ... C++ and parallel programming (e.g., CUDA). • Hands-on experience with large-scale model training ...
Intern, AI Engineering
San Francisco, CA · On-site
$19.75 - $25.50/hr
... Interns in San Francisco. You'll work on fundamental problems in LLM-based agentic systems and ... Proficiency in CUDA programming and custom kernel development for LLM operations * Background in ...
Intern, AI Engineering
San Francisco, CA · On-site
$19.75 - $25.50/hr
... Interns in San Francisco. You'll work on fundamental problems in LLM-based agentic systems and ... Proficiency in CUDA programming and custom kernel development for LLM operations * Background in ...
Internship Cuda Programmer information
What is the difference between Internship Cuda Programmer vs Cuda Developer?
| Aspect | Internship Cuda Programmer | Cuda Developer |
|---|---|---|
| Credentials | Enrolled in or recent graduate of Computer Science or related field | Bachelor's or higher in Computer Science, with experience in CUDA programming |
| Work Environment | Internship setting, learning-focused, entry-level projects | Full-time professional role, developing complex GPU-accelerated applications |
| Industry Usage | Research labs, tech companies, internships for skill development | Tech firms, gaming, scientific computing, high-performance computing |
While an Internship Cuda Programmer is typically a learning position for students or recent graduates gaining foundational experience, a Cuda Developer is a full-time professional responsible for designing and optimizing GPU-accelerated software. The roles differ mainly in experience level, responsibilities, and career stage, but both require knowledge of CUDA programming and GPU architecture.
- Internship Xr Software Engineer
- Software Engineer Intern
- Internship Master Software Engineer
- Internship Orbital Analyst
- Summer Software Engineer Internship
- Intern Compiler Engineer
- Commission Yahoo Software Engineer
- C++ Software Engineer Internship
- Google Cloud Platform Internship
- Internship 2024 Software Engineer New Grad
Inference Optimization Intern - Performance Modeling
Sunnyvale, CA • On-site
Internship
Posted 4 days ago
Job description
The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting-edge foundation models while pushing the limits of high-performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real-world impact.
As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures. This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.
Key Responsibilities
This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.
Responsibilities include:
- Develop analytical performance models for GPU kernels and inference workloads.
- Build and validate a simulator to estimate theoretical hardware performance limits.
- Compare measured kernel performance against architectural peak throughput.
- Identify performance bottlenecks in compute, memory, communication, and scheduling.
- Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
- Investigate PTX and SASS code generation to understand low-level execution behavior.
- Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
- Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
- Design profiling methodologies for Hopper and Blackwell architectures.
- Document findings and provide actionable recommendations for performance improvements.
Academic Qualifications
Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
Preferred Qualifications
- Experience with CUDA programming and GPU kernel development.
- Understanding of NVIDIA GPU architecture and memory hierarchy.
- Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
- Knowledge of PTX, SASS, and low-level GPU execution.
- Experience optimizing CUDA kernels for throughput and latency.
- Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
- Experience with deep learning frameworks such as PyTorch or TensorFlow.
- Strong programming skills in C++, CUDA, and Python.
Desired Skills
- Performance engineering mindset.
- Strong analytical and debugging abilities.
- Interest in AI systems, inference optimization, and hardware-software co-design.
- Ability to work independently on research and engineering challenges.
- Excellent written and verbal communication skills.