NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:

60 Nvidia Site Reliability Engineer Jobs Hiring Near You
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:
Senior Site Reliability Engineer - Storage
Santa Clara, CA · On-site
$168 - $322/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a ...
Senior Site Reliability Engineer - Storage
Santa Clara, CA · On-site
$168 - $322/hr
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Join our team at NVIDIA as a Senior Site reliability engineer focused on HPC storage and play a ...
Senior Site Reliability Engineer, AIOPs
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:
Senior Site Reliability Engineer, AIOPs
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... SRE/DevOps/Platform Ops. * Proven ownership of reliability for an observability/AIOps platform:
WI · On-site
$168 - $334/hr
At NVIDIA, our Compute Infrastructure Support (CIS) team looks for a driven Site Reliability Engineer passionate about innovation and excellence. As an SRE, you will be vital in delivering world ...
WI · On-site
$168 - $334/hr
At NVIDIA, our Compute Infrastructure Support (CIS) team looks for a driven Site Reliability Engineer passionate about innovation and excellence. As an SRE, you will be vital in delivering world ...
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Director, Engineering Operations and Site Reliability Engineering - Datacenter Server Systems
Santa Clara, CA · On-site
$67 - $89/hr
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than ... NVIDIA is seeking a strong technology leader for our Engineering Operations and Site Reliability ...
Senior Systems Software Engineer, Observability and Telemetry Platform
Santa Clara, CA · On-site
$184 - $287.50/hr
Senior Systems Software Engineer (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination ...
Senior Systems Software Engineer, Observability and Telemetry Platform
Santa Clara, CA · On-site
$184 - $287.50/hr
Senior Systems Software Engineer (SRE) at NVIDIA is an engineering discipline to design, build and maintain large scale production systems with high efficiency and availability using the combination ...
Senior Director, Enterprise Networking
Santa Clara, CA · Hybrid
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
Senior Director, Enterprise Networking
Santa Clara, CA · Hybrid
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
Senior Production Engineer - DGX Cloud
$57 - $76.25/hr
NVIDIA is hiring experienced Senior Production Engineers to help scale up its AI Infrastructure. We expect you to have significant experience with site reliability principles and techniques including ...
Senior Production Engineer - DGX Cloud
$57 - $76.25/hr
NVIDIA is hiring experienced Senior Production Engineers to help scale up its AI Infrastructure. We expect you to have significant experience with site reliability principles and techniques including ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
Senior Director, Enterprise Networking
Santa Clara, CA · On-site
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
Senior Director, Enterprise Networking
Santa Clara, CA · On-site
$127K - $173K/yr
Establish agentic, data-driven, SRE-oriented operating model with well-defined SLOs, observability ... NVIDIA is widely considered to be one of the technology world's most desirable employers. We have ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 ... We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering ...
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 ... We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering ...
NVIDIA is now passionate about innovation at the intersection of visual processing, high ... Reliability Engineer. The position is an individual contributor role, working with the cross ...
NVIDIA is now passionate about innovation at the intersection of visual processing, high ... Reliability Engineer. The position is an individual contributor role, working with the cross ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High ... Build and maintain SRE-style validation infrastructure, including provisioning, monitoring, and ...
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 ... We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering ...
NVIDIA has continuously reinvented itself over two decades. Our invention of the GPU in 1999 ... We are looking for a System Reliability Engineer to join NVIDIA's existing Reliability Engineering ...
Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational ... NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive ...
Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational ... NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive ...
Join NVIDIA, an innovator in computer graphics, PC gaming, and accelerated computing, as we step into the next era shaped by AI. As a Senior Reliability Engineer, you'll work within a focused team ...
Join NVIDIA, an innovator in computer graphics, PC gaming, and accelerated computing, as we step into the next era shaped by AI. As a Senior Reliability Engineer, you'll work within a focused team ...
NVIDIA is now passionate about innovation at the intersection of visual processing, high ... Reliability Engineer. The position is an individual contributor role, working with the cross ...
NVIDIA is now passionate about innovation at the intersection of visual processing, high ... Reliability Engineer. The position is an individual contributor role, working with the cross ...
Senior Engineer System Software, SDN Operations
Santa Clara, CA · On-site
$70.50 - $91.50/hr
Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational ... NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive ...
Senior Engineer System Software, SDN Operations
Santa Clara, CA · On-site
$70.50 - $91.50/hr
Collaborate with SRE, DevOps, and network engineering teams on production readiness and operational ... NVIDIA uses AI tools in its recruiting processes. NVIDIA is committed to fostering an inclusive ...
Nvidia Jobs Information
What is it like to work at Nvidia?
What makes Nvidia an attractive place to work?
Do workers at Nvidia get paid breaks?
45% of people say they don’t get paid breaks.
Based on data from 11 people who took the Breakroom Quiz between December 2024 and June 2026.
Does Nvidia pay people when they’re sick?
75% of people say they would get paid if they were sick but scheduled to work.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.
At Nvidia, are sick days and vacation days separate paid time off?
92% of people say they don’t have to use vacation days when they’re out sick.
Based on data from 13 people who took the Breakroom Quiz between June 2025 and June 2026.
Is the health insurance from Nvidia affordable enough for their workers?
100% of people say the health insurance costs are okay
Based on data from 12 people who took the Breakroom Quiz between June 2025 and June 2026.
Do people get paid time off at Nvidia?
100% of people say they get paid time off.
Based on data from 13 people who took the Breakroom Quiz between June 2025 and June 2026.
How far ahead of time do people find out their work schedule?
- 67% of people with changing schedules find out their shifts one week or less ahead of time.
- 0% of people with changing schedules find out their shifts two weeks ahead of time.
- 0% of people with changing schedules find out their shifts three weeks ahead of time.
- 33% of people with changing schedules find out their shifts four weeks or more ahead of time.
Based on data from 6 people who took the Breakroom Quiz between April 2025 and March 2026.
Do workers at Nvidia worry about hours?
83% of people report they don’t worry about getting enough hours.
Based on data from 12 people who took the Breakroom Quiz between December 2024 and March 2026.
Do Nvidia workers get to choose the shifts they work?
40% report that they don’t have enough control over which shifts they work.
Based on data from 5 people who took the Breakroom Quiz between December 2024 and January 2026.
How easy is it for Nvidia workers to change shifts?
88% of people report that it’s easy to change shifts if they need to.
Based on data from 8 people who took the Breakroom Quiz between January 2025 and March 2026.
How easy is it to get time off at Nvidia?
100% of people report it’s easy to get time off.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do Nvidia managers change schedules at the last minute?
92% of people say their manager doesn’t change their shift schedule at the last minute.
Based on data from 13 people who took the Breakroom Quiz between December 2024 and March 2026.
Do jobs at Nvidia spill into time workers aren’t paid for?
17% of people report that their job takes up time that they don’t get paid for.
Based on data from 12 people who took the Breakroom Quiz between December 2024 and March 2026.
How easy is it to take sick days at Nvidia?
100% of people report that it’s easy to take time off if they are sick.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia feel treated with respect by their managers?
100% of people say they’re treated with respect by their managers.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia get to take their breaks without interruption?
94% of people report that they get to take their breaks without interruption.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Is it stressful to work at Nvidia?
47% of people say they often feel stressed at work.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia enjoy their jobs?
100% of people report they enjoy their job.
Based on data from 15 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people at Nvidia recommend working with their team?
76% of people report that they would recommend working with their immediate team to a friend.
Based on data from 17 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people get enough training when they start at Nvidia?
38% of people report they didn’t get enough training when they started working here.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people get support to advance at Nvidia?
In the last year, 93% of people report being given support to advance their career here.
Based on data from 15 people who took the Breakroom Quiz between December 2024 and June 2026.
Do people think Nvidia’s headquarters understands what’s happening where they work?
67% of people think that this employer’s headquarters or owners have a good understanding of what’s really happening where they work.
Based on data from 15 people who took the Breakroom Quiz between December 2024 and June 2026.
Do workers feel well informed about how Nvidia is doing?
94% of people feel that they are kept well informed about how the company is doing as a whole.
Based on data from 16 people who took the Breakroom Quiz between December 2024 and June 2026.
What other companies are hiring for Site Reliability Engineer jobs?
What are the most popular jobs at Nvidia?
What are the most popular categories at Nvidia?

$67 - $89/hr
Full-time
Re-posted 26 days ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
7th of 245 rated software companies
Job description
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
Join our team of innovative engineers who are building an AI Data Center AIOps platform that turns raw, high-volume telemetry into reliable, job-centric insights and automation for GPU fleets. We're hiring a DevOps Engineer to operate the platform itself (not the compute cluster): uptime, performance, data integrity, and safe change management. You'll own SLOs/SLIs, incident response, and postmortems for the telemetry ingestion, processing, storage, and APIs/dashboards that operators depend on. You'll partner Software Engineering and Systems Engineering team to translate platform signals into actionable, trustworthy alerts and automation.
What you'll be doing:
Continuously monitor platform health via dashboards/logs/metrics, automate recurring checks, and keep reliability + resource efficiency on track.
Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation), and lead rollbacks/remediations when needed.
Lead first-level incident triage: collect diagnostics, identify likely root causes, and hand off clear, actionable findings to engineering.
Build and maintain runbooks/SOPs/checklists, pushing continuous improvement through automation.
Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.
Contribute in adjacent functional areas to grow and help your team members!
What we need to see:
BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.
Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.
Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry-heavy microservices-ingestion, processing, storage, APIs, and UI.
Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.
Clear communicator who writes excellent runbooks/docs and can translate ambiguous requirements into concrete operational practices and dependable customer-facing reliability.
Ways to stand out from the crowd:
Strong Linux + networking fundamentals, distributed systems instincts, and hands-on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.
Experience building safe automation that operators trust: canary releases, automated rollback criteria, "monitoring for the monitoring" (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.
Strong in distributed/streaming systems operations (Kafka/Pulsar, Flink/Spark, ClickHouse/Elastic/TSDBs, object storage)-and can reason about backpressure, hotspots, and failure domains end-to-end.
Proven programming experience building automation tools or services - ideally in Python, or similar languages - to simplify operations and scale recurring processes.
Proven experience running largescale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with handson experience with observability tools - you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.
With competitive salaries and a generous benefits package, we are widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and hardworking people in the world working for us and, due to unprecedented growth, our exclusive engineering teams are rapidly growing. If you're a creative and autonomous engineer with a real passion for technology, we want to hear from you.
Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 148,000 USD - 235,750 USD for Level 3, and 176,000 USD - 276,000 USD for Level 4.You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993