You'll architect and implement high-performance inference software, optimize GPU kernels, drive ... Strong programming skills in Python and one of C/C++, Go, or Rust. Solid CS fundamentals ...
You'll architect and implement high-performance inference software, optimize GPU kernels, drive ... Strong programming skills in Python and one of C/C++, Go, or Rust. Solid CS fundamentals ...
Understanding of GPU architecture and system-level performance concepts * Experience with ... Experience with AI-powered developer tools and ability to effectively leverage them in day-to-day ...
Understanding of GPU architecture and system-level performance concepts * Experience with ... Experience with AI-powered developer tools and ability to effectively leverage them in day-to-day ...
Familiarity with GPU performance analysis: memory hierarchy, utilization, roofline modeling, and ... Experience building developer tools or internal platforms that meaningfully improved team ...
Familiarity with GPU performance analysis: memory hierarchy, utilization, roofline modeling, and ... Experience building developer tools or internal platforms that meaningfully improved team ...
Understanding of GPU architecture and system-level performance concepts * Experience with ... Experience with AI-powered developer tools and ability to effectively leverage them in day-to-day ...
Understanding of GPU architecture and system-level performance concepts * Experience with ... Experience with AI-powered developer tools and ability to effectively leverage them in day-to-day ...
Senior / Staff Graphics Software Engineer
Toronto, ON · On-site +1
CA$155K - CA$269K/yr
You have a deep understanding of GPU performance, and you know how to use instrumentation and ... engineering fundamentals. You write efficient and maintainable code in systems languages such as C ...
Senior / Staff Graphics Software Engineer
Toronto, ON · On-site +1
CA$155K - CA$269K/yr
You have a deep understanding of GPU performance, and you know how to use instrumentation and ... engineering fundamentals. You write efficient and maintainable code in systems languages such as C ...
... Engineers to verify the design and implementation of the world's leading SoC's and GPU's. This ... In this position, you will help to build the high-performance processor elements that implement ...
... Engineers to verify the design and implementation of the world's leading SoC's and GPU's. This ... In this position, you will help to build the high-performance processor elements that implement ...
AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI ... As an Lead SMU (System Management Unit) Validation Engineer in the Data Center and GPU Accelerated ...
AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI ... As an Lead SMU (System Management Unit) Validation Engineer in the Data Center and GPU Accelerated ...
AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI ... As an Lead SMU (System Management Unit) Validation Engineer in the Data Center and GPU Accelerated ...
AMD is powering the next generation of supercomputing, high-performance computing, cloud, and AI ... As an Lead SMU (System Management Unit) Validation Engineer in the Data Center and GPU Accelerated ...
CPU Performance Architect
Markham, ON · On-site
The successful candidate will work in AMD's Client and Graphics SOC Performance Team in Markham ... CPU, GPU, NPU, Video Processors, and/or DRAM controllers * Proficient in C/C++ programming and ...
CPU Performance Architect
Markham, ON · On-site
The successful candidate will work in AMD's Client and Graphics SOC Performance Team in Markham ... CPU, GPU, NPU, Video Processors, and/or DRAM controllers * Proficient in C/C++ programming and ...
We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm ... GPU software stacks (ROCm, CUDA, oneAPI, SYCL) * AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton ...
We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm ... GPU software stacks (ROCm, CUDA, oneAPI, SYCL) * AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton ...
Principal Software Developer - AI/ML Performance Validation & Systems Testing
Thornhill, ON · On-site
We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm ... GPU software stacks (ROCm, CUDA, oneAPI, SYCL) * AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton ...
Principal Software Developer - AI/ML Performance Validation & Systems Testing
Thornhill, ON · On-site
We are seeking a Principal Software Engineer to serve as the senior technical leader for ROCm ... GPU software stacks (ROCm, CUDA, oneAPI, SYCL) * AI/ML frameworks (PyTorch, TensorFlow, JAX, Triton ...
Senior Platform Engineer
Toronto, ON · On-site
Quantiphi is an award-winning, AI-First digital engineering and consulting company focused on ... Perform GPU profiling, benchmarking, and performance optimization for distributed training ...
Senior Platform Engineer
Toronto, ON · On-site
Quantiphi is an award-winning, AI-First digital engineering and consulting company focused on ... Perform GPU profiling, benchmarking, and performance optimization for distributed training ...
CPU Performance Architect
Markham, ON · On-site
The successful candidate will work in AMD's Client and Graphics SOC Performance Team in Markham ... CPU, GPU, NPU, Video Processors, and/or DRAM controllers * Proficient in C/C++ programming and ...
CPU Performance Architect
Markham, ON · On-site
The successful candidate will work in AMD's Client and Graphics SOC Performance Team in Markham ... CPU, GPU, NPU, Video Processors, and/or DRAM controllers * Proficient in C/C++ programming and ...
Write and optimize high-performance GPU kernels (GEMM, attention, quantized matmul, GPTQ/AWQ) in ... Contribute to the open ROCm ecosystem and AMD's developer experience -- SDKs, CI dashboards ...
Write and optimize high-performance GPU kernels (GEMM, attention, quantized matmul, GPTQ/AWQ) in ... Contribute to the open ROCm ecosystem and AMD's developer experience -- SDKs, CI dashboards ...
Site Reliability Engineer
Toronto, ON · On-site +1
CA$125K - CA$250K/yr
... performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps ... engineering standards Minimum Qualifications * 4+ years of experience in site reliability ...
Site Reliability Engineer
Toronto, ON · On-site +1
CA$125K - CA$250K/yr
... performance networks, GPU clusters, storage, scheduling, and the operational tooling that keeps ... engineering standards Minimum Qualifications * 4+ years of experience in site reliability ...
🚀 Staff LLMOps Engineer (Cloud / AI Infrastructure) Location: Downtown Toronto Hybrid: 4 days in ... GPU clusters and pushing the boundaries of inference performance for enterprise-grade AI ...
Quick apply
🚀 Staff LLMOps Engineer (Cloud / AI Infrastructure) Location: Downtown Toronto Hybrid: 4 days in ... GPU clusters and pushing the boundaries of inference performance for enterprise-grade AI ...
Principal Graphics Engineer
Toronto, ON · On-site
... ship performance-critical code that holds up at scale. We are equally interested in candidates with a deep graphics engineering background (game engines, browser graphics, GPU systems, real-time ...
Quick apply
Principal Graphics Engineer
Toronto, ON · On-site
... ship performance-critical code that holds up at scale. We are equally interested in candidates with a deep graphics engineering background (game engines, browser graphics, GPU systems, real-time ...
Principal Graphics Engineer
Toronto, ON · On-site
... ship performance-critical code that holds up at scale. We are equally interested in candidates with a deep graphics engineering background (game engines, browser graphics, GPU systems, real-time ...
Principal Graphics Engineer
Toronto, ON · On-site
... ship performance-critical code that holds up at scale. We are equally interested in candidates with a deep graphics engineering background (game engines, browser graphics, GPU systems, real-time ...
We are seeking a Senior Board Hardware Debug Engineer to join our Datacenter & AI Board Validation ... Own complex board-level debug for high-performance Datacenter GPU platforms * Drive hardware ...
We are seeking a Senior Board Hardware Debug Engineer to join our Datacenter & AI Board Validation ... Own complex board-level debug for high-performance Datacenter GPU platforms * Drive hardware ...
We are seeking a Senior Board Hardware Debug Engineer to join our Datacenter & AI Board Validation ... Own complex board-level debug for high-performance Datacenter GPU platforms * Drive hardware ...
We are seeking a Senior Board Hardware Debug Engineer to join our Datacenter & AI Board Validation ... Own complex board-level debug for high-performance Datacenter GPU platforms * Drive hardware ...
Gpu Performance Engineer information
What are some common challenges faced by a GPU performance engineer when optimizing graphics workloads?
What is a GPU performance engineer?
What are the key skills and qualifications needed to thrive as a GPU performance engineer, and why are they important?
What is the difference between Gpu Performance Engineer vs Gpu Hardware Engineer?
| Aspect | Gpu Performance Engineer | Gpu Hardware Engineer |
|---|---|---|
| Primary Focus | Optimizing GPU performance, benchmarking, and tuning software | Designing, developing, and testing GPU hardware components |
| Required Skills | Programming, performance analysis, GPU architecture knowledge | Hardware design, circuit analysis, FPGA/ASIC experience |
| Work Environment | Software development teams, labs for testing performance | Hardware labs, manufacturing facilities, R&D centers |
| Common Certifications | None specific, often requires computer engineering or related degrees | Electrical engineering, VLSI design certifications |
The Gpu Performance Engineer primarily focuses on optimizing and testing GPU software performance, while the Gpu Hardware Engineer designs and develops the physical GPU components. Both roles require a strong background in computer engineering, but differ in their core responsibilities and work environments.

Full-time
Posted 4 days ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
8th of 242 rated software companies
Job description
We are seeking highly skilled and motivated software engineers to join us and build AI inference systems that serve large-scale models with extreme efficiency. You'll architect and implement high-performance inference software, optimize GPU kernels, drive industry benchmarks, and work with state-of-the-art research techniques to improve serving efficiency. You'll collaborate across inference performance, kernels, training, large-scale serving, and research teams to push the frontier of accelerated computing for AI.
What you'll be doing:
Contribute features to vLLM that empower the newest models with the latest NVIDIA GPU hardware features and serving runtime algorithms.
Profile and optimize the inference framework (vLLM) with methods like speculative decoding, 5D Parallelism, and prefill-decode disaggregation.
Architect novel frameworks and runtime optimizations for inference infrastructure, benchmarking, and kernels.
Conduct and publish original research that advances the Pareto frontier in ML Systems; survey recent publications and find a way to integrate research ideas and prototypes into production-grade, open-source software.
Develop, optimize, and benchmark GPU kernels (both hand-tuned and compiler-generated) using techniques such as fusion, autotuning, and memory/layout optimization.
What we need to see:
Bachelor's, Master's, or PhD degree in Computer Science (CS), Computer Engineering (CE) or Software Engineering (SE).
5+ years of industry experience in software engineering or equivalent research experience.
Strong programming skills in Python and one of C/C++, Go, or Rust. Solid CS fundamentals: algorithms & data structures, operating systems, computer architecture, parallel programming, software engineering, distributed systems, deep learning theories.
Knowledgeable and passionate about performance engineering in ML frameworks (e.g., PyTorch) and inference engines (e.g., vLLM and SGLang).
Familiarity with GPU programming and performance: CUDA, memory hierarchy, streams, NCCL; proficiency with profiling/debug tools (e.g., Nsight Systems/Compute).
Excellent debugging, problem-solving, and communication skills; ability to excel in a fast-paced, multi-functional setting.
Ways to Stand out from the Crowd
Experience developing major features and optimizations for LLM inference engines (e.g., vLLM, SGLang).
Hands-on work with LLM inference and training runtimes (deploying LLMs to production, large-scale LLM pre-training and RL), ML compilers and DSLs (e.g., Triton, CuTe, MLIR/LLVM, XLA), GPU libraries (e.g., CUTLASS) and features (e.g., CUDA Graph, Tensor Cores).
Experience with speculative decoding training and runtime features: tree-structured drafting, parallel drafting, diffusion LLMs, DFlash, EAGLE.
Contributions to open-source projects and/or publications; please include links to GitHub pull requests, published papers and artifacts.
At NVIDIA, we believe artificial intelligence (AI) will fundamentally transform how people live and work. Our mission is to advance AI research and development to create groundbreaking technologies that enable anyone to harness the power of AI and benefit from its potential.
Our team consists of experts in AI, systems and performance optimization. Our leadership includes world-renowned experts in AI systems who have received multiple academic and industry research awards. If you're excited to build systems, kernels, and tools that make large-scale AI faster, more efficient, and easier to deploy, we'd love to hear from you.
#LI-Hybrid
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993