Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute. * Investigate PTX and SASS code generation to understand low-level execution behavior. * Collaborate with researchers and ...
Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute. * Investigate PTX and SASS code generation to understand low-level execution behavior. * Collaborate with researchers and ...
Senior Research Engineer - AI Coding Tools
Santa Clara, CA · On-site
$143K - $189K/yr
We make NVIDIA's core developer tools - including Nsight Compute and Nsight Systems - first-class citizens for AI agents through MCP servers and Agent Skills. The space shifts every few weeks, and we ...
Senior Research Engineer - AI Coding Tools
Santa Clara, CA · On-site
$143K - $189K/yr
We make NVIDIA's core developer tools - including Nsight Compute and Nsight Systems - first-class citizens for AI agents through MCP servers and Agent Skills. The space shifts every few weeks, and we ...
Senior Research Engineer - AI Coding Tools
Santa Clara, CA · On-site
$143K - $189K/yr
We make NVIDIA's core developer tools - including Nsight Compute and Nsight Systems - first-class citizens for AI agents through MCP servers and Agent Skills. The space shifts every few weeks, and we ...
Senior Research Engineer - AI Coding Tools
Santa Clara, CA · On-site
$143K - $189K/yr
We make NVIDIA's core developer tools - including Nsight Compute and Nsight Systems - first-class citizens for AI agents through MCP servers and Agent Skills. The space shifts every few weeks, and we ...
Senior Research Engineer - AI Coding Tools
Santa Clara, CA · On-site
$143K - $189K/yr
... tools (Nsight Compute, Nsight Systems, and more) first-class for AI agents • Generate, curate, and validate synthetic training and evaluation data for CUDA programming • Deliver "net new ...
Senior Research Engineer - AI Coding Tools
Santa Clara, CA · On-site
$143K - $189K/yr
... tools (Nsight Compute, Nsight Systems, and more) first-class for AI agents • Generate, curate, and validate synthetic training and evaluation data for CUDA programming • Deliver "net new ...
Software Engineer - Kernels/CUDA (C++)
Palo Alto, CA · On-site
$180K - $440K/yr
Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance * Profile, debug, and eliminate bottlenecks across GPU memory ...
Software Engineer - Kernels/CUDA (C++)
Palo Alto, CA · On-site
$180K - $440K/yr
Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance * Profile, debug, and eliminate bottlenecks across GPU memory ...
... Nsight for maximum performance • Work on Linux kernel internals, scheduling, memory management, and resource isolation at cluster scale • Build custom container orchestration, virtualization ...
... Nsight for maximum performance • Work on Linux kernel internals, scheduling, memory management, and resource isolation at cluster scale • Build custom container orchestration, virtualization ...
You use tools like Nsight Systems / Nsight Compute to find bottlenecks, validate hypotheses, and iterate until improvements show up in end-to-end benchmarks. * Bridges theory and practice: You can ...
You use tools like Nsight Systems / Nsight Compute to find bottlenecks, validate hypotheses, and iterate until improvements show up in end-to-end benchmarks. * Bridges theory and practice: You can ...
Staff ML Performance Engineer (Training Efficiency)
$336K - $359K/yr
Profile ML workloads to identify their bottlenecks, e.g. using NVIDIA Nsight Systems * Design and implement efficiency improvements to maximize MFU and throughput, e.g. parallelism, model compilation ...
Staff ML Performance Engineer (Training Efficiency)
$336K - $359K/yr
Profile ML workloads to identify their bottlenecks, e.g. using NVIDIA Nsight Systems * Design and implement efficiency improvements to maximize MFU and throughput, e.g. parallelism, model compilation ...
Staff ML Performance Engineer (Training Efficiency)
Sunnyvale, CA · On-site
$336K - $359K/yr
Profile ML workloads to identify their bottlenecks, e.g. using NVIDIA Nsight Systems * Design and implement efficiency improvements to maximize MFU and throughput, e.g. parallelism, model compilation ...
Staff ML Performance Engineer (Training Efficiency)
Sunnyvale, CA · On-site
$336K - $359K/yr
Profile ML workloads to identify their bottlenecks, e.g. using NVIDIA Nsight Systems * Design and implement efficiency improvements to maximize MFU and throughput, e.g. parallelism, model compilation ...
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara, CA · On-site
$164K/yr
Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.
Inference Performance Engineer, AI Inference Configuration Optimization
Santa Clara, CA · On-site
$164K/yr
Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.
System Software Engineer, Robot Platform -- GPU & Accelerated Compute
Redwood City, CA · On-site
$211K - $251K/yr
CUDA runtime API, CUDA Graphs, and CUDA IPC • Familiarity with GPU sharing mechanisms such as MPS and MIG • Experience with GPU profiling tools such as Nsight Systems and Nsight Compute • Solid ...
System Software Engineer, Robot Platform -- GPU & Accelerated Compute
Redwood City, CA · On-site
$211K - $251K/yr
CUDA runtime API, CUDA Graphs, and CUDA IPC • Familiarity with GPU sharing mechanisms such as MPS and MIG • Experience with GPU profiling tools such as Nsight Systems and Nsight Compute • Solid ...
System Software Engineer -- GPU & Accelerated Compute
Redwood City, CA · On-site
$211K - $251K/yr
CUDA runtime API, CUDA Graphs, and CUDA IPC • Familiarity with GPU sharing mechanisms such as MPS and MIG • Experience with GPU profiling tools such as Nsight Systems and Nsight Compute • Solid ...
System Software Engineer -- GPU & Accelerated Compute
Redwood City, CA · On-site
$211K - $251K/yr
CUDA runtime API, CUDA Graphs, and CUDA IPC • Familiarity with GPU sharing mechanisms such as MPS and MIG • Experience with GPU profiling tools such as Nsight Systems and Nsight Compute • Solid ...
Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.
Profile workloads using Nsight Systems, kernel traces, and internal analysis tools. Use roofline and speed-of-light analysis to find credible headroom and drive fixes from hypothesis to measured wins.
AI/ML Infrastructure Engineer
San Francisco, CA · On-site
$126K - $166K/yr
Experience with profiling and benchmarking tools (e.g., Nsight Systems, Nsight Compute) to validate performance on complex architectures. * Experience identifying and resolving compute and data flow ...
Quick apply
AI/ML Infrastructure Engineer
San Francisco, CA · On-site
$126K - $166K/yr
Experience with profiling and benchmarking tools (e.g., Nsight Systems, Nsight Compute) to validate performance on complex architectures. * Experience identifying and resolving compute and data flow ...
Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation * Write high-performance CUDA and Triton kernels for critical model operations * Optimize cold start ...
Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation * Write high-performance CUDA and Triton kernels for critical model operations * Optimize cold start ...
Senior AI Performance and Efficiency Engineer
Santa Clara, CA · On-site
$122K - $168K/yr
... NSight Compute • Experience with debugging large-scale distributed training using NCCL • Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with ...
Senior AI Performance and Efficiency Engineer
Santa Clara, CA · On-site
$122K - $168K/yr
... NSight Compute • Experience with debugging large-scale distributed training using NCCL • Proficiency in programming & scripting languages such as Python, Go, Bash, as well as familiarity with ...
... Nsight Compute, Nsight Systems, nvprof, or equivalent • Solid understanding of GPU architecture -- SMs, memory hierarchy, power states, and how they map to current draw profiles • Working ...
... Nsight Compute, Nsight Systems, nvprof, or equivalent • Solid understanding of GPU architecture -- SMs, memory hierarchy, power states, and how they map to current draw profiles • Working ...
Nsight Compute, Nsight Systems, rocprof, or equivalents • Experience reading and reasoning about PTX/SASS or GPU assembly • Solid systems programming in C++ and CUDA (or ROCm/HIP) • Good ...
Nsight Compute, Nsight Systems, rocprof, or equivalents • Experience reading and reasoning about PTX/SASS or GPU assembly • Solid systems programming in C++ and CUDA (or ROCm/HIP) • Good ...
... Nsight Compute, Nsight Systems, nvprof, or equivalent • Solid understanding of GPU architecture -- SMs, memory hierarchy, power states, and how they map to current draw profiles • Working ...
... Nsight Compute, Nsight Systems, nvprof, or equivalent • Solid understanding of GPU architecture -- SMs, memory hierarchy, power states, and how they map to current draw profiles • Working ...
Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation * Write high-performance CUDA and Triton kernels for critical model operations * Optimize cold start ...
Quick apply
Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation * Write high-performance CUDA and Triton kernels for critical model operations * Optimize cold start ...
Nsight information
What types of projects do employees typically work on at Nsight, and how do teams collaborate to achieve project goals?
What are Nsight jobs?
What are the key skills and qualifications needed to thrive as a data analyst at Nsight?
What is the difference between Nsight vs Network Security Analyst?
| Aspect | Nsight | Network Security Analyst |
|---|---|---|
| Required Certifications | Typically Cisco, CompTIA Security+ | CompTIA Security+, CISSP, CEH |
| Work Environment | IT consulting firms, tech companies, remote options | Corporate IT departments, security firms, government agencies |
| Industry Usage | Technology, consulting, cybersecurity | Cybersecurity, IT, finance, government |
| Common Search/Comparison | Yes | Yes |
While Nsight professionals focus on providing IT consulting and solutions, Network Security Analysts specialize in protecting networks from threats. Both roles require cybersecurity certifications and work in tech environments, but Nsight roles often involve broader IT consulting, whereas Network Security Analysts focus specifically on security measures and threat mitigation.

Inference Optimization Intern - Performance Modeling
Sunnyvale, CA • On-site
Internship
Re-posted 17 days ago
Job description
The Institute of Foundation Models is dedicated to advancing the science and engineering of large-scale AI systems. Our researchers and engineers develop cutting-edge foundation models while pushing the limits of high-performance computing and efficient AI inference. By combining deep expertise in machine learning, systems engineering, and hardware optimization, we build scalable AI solutions that drive scientific discovery and real-world impact.
As part of the team, interns work alongside world-class researchers and performance engineers to optimize the execution of large-scale foundation models on next-generation NVIDIA GPU architectures. This internship provides hands-on experience in low-level GPU performance analysis, kernel optimization, and hardware-aware inference acceleration.
Key Responsibilities
This intensive internship offers a unique opportunity to contribute to the development of a simulator and profiling framework for foundation model inference on NVidia GPUs.
Responsibilities include:
- Develop analytical performance models for GPU kernels and inference workloads.
- Build and validate a simulator to estimate theoretical hardware performance limits.
- Compare measured kernel performance against architectural peak throughput.
- Identify performance bottlenecks in compute, memory, communication, and scheduling.
- Analyze GPU execution using NVIDIA Nsight Systems and Nsight Compute.
- Investigate PTX and SASS code generation to understand low-level execution behavior.
- Collaborate with researchers and engineers to optimize inference kernels for transformer-based models.
- Evaluate utilization of Tensor Cores, memory bandwidth, caches, and instruction pipelines.
- Design profiling methodologies for Hopper and Blackwell architectures.
- Document findings and provide actionable recommendations for performance improvements.
Academic Qualifications
Currently pursuing a degree in Computer Science, Computer Engineering, Electrical Engineering, Artificial Intelligence, High-Performance Computing, or a related quantitative discipline.
Preferred Qualifications
- Experience with CUDA programming and GPU kernel development.
- Understanding of NVIDIA GPU architecture and memory hierarchy.
- Familiarity with performance profiling tools such as Nsight Systems and Nsight Compute.
- Knowledge of PTX, SASS, and low-level GPU execution.
- Experience optimizing CUDA kernels for throughput and latency.
- Understanding of roofline analysis, performance modeling, and hardware utilization metrics.
- Experience with deep learning frameworks such as PyTorch or TensorFlow.
- Strong programming skills in C++, CUDA, and Python.
Desired Skills
- Performance engineering mindset.
- Strong analytical and debugging abilities.
- Interest in AI systems, inference optimization, and hardware-software co-design.
- Ability to work independently on research and engineering challenges.
- Excellent written and verbal communication skills.