At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges ...

20 Nvidia Site Reliability Engineer Jobs Hiring Near You
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges ...
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, CA · On-site
$67 - $89/hr
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges ...
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, CA · On-site
$67 - $89/hr
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges ...
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA is a leading technology company specializing in AI and computing solutions. They are seeking a Site Reliability Engineer to develop and support large-scale production systems, ensuring high ...
Site Reliability Engineer - Hardware Infrastructure
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA is a leading technology company specializing in AI and computing solutions. They are seeking a Site Reliability Engineer to develop and support large-scale production systems, ensuring high ...
Senior Site Reliability Engineer - HPC
Durham, NC · On-site
$55 - $73.25/hr
NVIDIA has been transforming computer graphics and accelerated computing for over 25 years. They are seeking a Senior Site Reliability Engineer to join their Compute Farm team, responsible for ...
Senior Site Reliability Engineer - HPC
Durham, NC · On-site
$55 - $73.25/hr
NVIDIA has been transforming computer graphics and accelerated computing for over 25 years. They are seeking a Senior Site Reliability Engineer to join their Compute Farm team, responsible for ...
Senior Site Reliability Engineer - HPC
Durham, NC · On-site
$55 - $73.25/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of ...
Senior Site Reliability Engineer - HPC
Durham, NC · On-site
$55 - $73.25/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of ...
Senior Site Reliability Engineer - HPC
$56.50 - $75/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of ...
Senior Site Reliability Engineer - HPC
$56.50 - $75/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... We're looking for a Senior SRE to join our Compute Farm team and help build the next generation of ...
Senior Site Reliability Engineer - HPC
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Own SRE solutions end-to-end, from design and implementation to operation and continuous ...
Senior Site Reliability Engineer - HPC
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Own SRE solutions end-to-end, from design and implementation to operation and continuous ...
Senior Site Reliability Engineer, AIOPs
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics and computing for over 25 years, and they are seeking a Senior Site Reliability Engineer to join their innovative team. The role involves operating an ...
Senior Site Reliability Engineer, AIOPs
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics and computing for over 25 years, and they are seeking a Senior Site Reliability Engineer to join their innovative team. The role involves operating an ...
Senior SRE Engineer
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... What we need to see: * 5+ years in SRE/platform roles with strong fundamentals - SLO/SLI build ...
Senior SRE Engineer
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... What we need to see: * 5+ years in SRE/platform roles with strong fundamentals - SLO/SLI build ...
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:
Senior Site Reliability Engineer, AIOPs
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:
Senior Site Reliability Engineer, AIOPs
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Senior Director, Enterprise Networking
Santa Clara, CA · Hybrid
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
Senior Director, Enterprise Networking
Santa Clara, CA · Hybrid
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
Demonstrated experience implementing SRE practices, specifically defining and tracking SLIs, SLOs ... NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive ...
Demonstrated experience implementing SRE practices, specifically defining and tracking SLIs, SLOs ... NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive ...
Senior Director, Enterprise Networking
Santa Clara, CA · On-site
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
Senior Director, Enterprise Networking
Santa Clara, CA · On-site
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 ... We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering ...
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 ... We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering ...
Nvidia Jobs Information
What is it like to work at Nvidia?
What makes Nvidia an attractive place to work?
Do workers at Nvidia get paid breaks?
45% of people say they don’t get paid breaks.
Based on data from 11 people who took the Breakroom Quiz between December 2024 and June 2026.
Does Nvidia pay people when they’re sick?
75% of people say they would get paid if they were sick but scheduled to work.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.
At Nvidia, are sick days and vacation days separate paid time off?
92% of people say they don’t have to use vacation days when they’re out sick.
Based on data from 13 people who took the Breakroom Quiz between June 2025 and June 2026.
Is the health insurance from Nvidia affordable enough for their workers?
100% of people say the health insurance costs are okay
Based on data from 12 people who took the Breakroom Quiz between June 2025 and June 2026.
Do people get paid time off at Nvidia?
100% of people say they get paid time off.
Based on data from 13 people who took the Breakroom Quiz between June 2025 and June 2026.
How far ahead of time do people find out their work schedule?
- 67% of people with changing schedules find out their shifts one week or less ahead of time.
- 0% of people with changing schedules find out their shifts two weeks ahead of time.
- 0% of people with changing schedules find out their shifts three weeks ahead of time.
- 33% of people with changing schedules find out their shifts four weeks or more ahead of time.
Based on data from 6 people who took the Breakroom Quiz between April 2025 and March 2026.
Do workers at Nvidia worry about hours?
83% of people report they don’t worry about getting enough hours.
Based on data from 12 people who took the Breakroom Quiz between December 2024 and March 2026.
Do Nvidia workers get to choose the shifts they work?
40% report that they don’t have enough control over which shifts they work.
Based on data from 5 people who took the Breakroom Quiz between December 2024 and January 2026.
How easy is it for Nvidia workers to change shifts?
88% of people report that it’s easy to change shifts if they need to.
Based on data from 8 people who took the Breakroom Quiz between January 2025 and March 2026.
How easy is it to get time off at Nvidia?
100% of people report it’s easy to get time off.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do Nvidia managers change schedules at the last minute?
92% of people say their manager doesn’t change their shift schedule at the last minute.
Based on data from 13 people who took the Breakroom Quiz between December 2024 and March 2026.
Do jobs at Nvidia spill into time workers aren’t paid for?
17% of people report that their job takes up time that they don’t get paid for.
Based on data from 12 people who took the Breakroom Quiz between December 2024 and March 2026.
How easy is it to take sick days at Nvidia?
100% of people report that it’s easy to take time off if they are sick.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia feel treated with respect by their managers?
100% of people say they’re treated with respect by their managers.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia get to take their breaks without interruption?
94% of people report that they get to take their breaks without interruption.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Is it stressful to work at Nvidia?
47% of people say they often feel stressed out at work.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia enjoy their jobs?
100% of people report they enjoy their job.
Based on data from 15 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia recommend working with their team?
76% of people report that they would recommend working with their immediate team to a friend.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people get enough training when they start at Nvidia?
38% of people report they didn’t get enough training when they started working here.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people get support to advance at Nvidia?
In the last year, 93% of people report being given support to advance their career here.
Based on data from 15 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people think Nvidia’s headquarters understands what’s happening where they work?
67% of people think that this employer’s headquarters or owners have a good understanding of what’s really happening where they work.
Based on data from 15 people who took the Breakroom Quiz between December 2024 and June 2026.
Do workers feel well informed about how Nvidia is doing?
94% of people feel that they are kept well informed about how the company is doing as a whole.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.

$67 - $89/hr
Full-time
Re-posted 15 days ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
8th of 246 rated software companies
Job description
At NVIDIA, Site Reliability Engineering provides a rare chance to define, develop, and support large-scale production systems with high efficiency and availability. This demanding position merges software and systems engineering efforts to guarantee flawless service operation with consistent reliability and uptime. As an SRE here, you will be part of a welcoming team that values collaboration and creativity, empowering developers to make significant updates while sustaining efficient system function.
What you'll be doing:
Develop and support guidelines for incident management, planned maintenance, and blameless postmortems.
Assist teams in responding to high severity incidents, driving root cause analysis, crafting high-quality postmortems, and developing post-incident corrective actions.
Define reliability and supportability metrics, Service Level Objectives, and error budgets.
Develop and drive the adoption of actionable, customer-centric monitoring and alerting.
Apply automation and Generative AI/Agentic solutions to minimize manual and tedious activities and boost customer support.
Guide teams on establishing sustainable on-call and operational standards.
What we need to see:
Degree in Computer Science or a related technical field involving coding, or equivalent experience.
8+ years of experience in SRE, DevOps, or Production Engineering.
Strong understanding of SRE principles, including incident management, error budgets, SLOs, and SLAs.
Experience crafting and deploying systems that are fault-tolerant, performant, and supportable.
Background with infrastructure automation.
Experience running critical services in production.
Experience in one or more of the following: Python, Go, Perl, or Ruby.
Hands-on experience with observability platforms (e.g., Prometheus, Grafana).
Strong communication skills with the ability to convey technical concepts effectively to diverse audiences.
Flexibility and adaptability working in a fast-paced environment with evolving requirements.
Ways to stand out from the crowd:
Expertise in establishing incident management and postmortem processes.
Experience driving adoption of common tools and processes across diverse groups.
Experience working with LLM/Generative AI/Agentic solutions to shorten mitigation time, lessen toil, and ensure Service Level Objectives are met.
Hands-on expertise operating and scaling distributed systems with tight SLAs, ensuring high availability and performance.
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993