$108K - $147K/yr
Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management. * Build observability and security capabilities that improve the reliability ...
$108K - $147K/yr
Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management. * Build observability and security capabilities that improve the reliability ...
$108K - $147K/yr
Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management. * Build observability and security capabilities that improve the reliability ...
$112K/yr
Experience with Infrastructure as Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Prometheus, Grafana, and ELK or ...
$112K/yr
Experience with Infrastructure as Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Prometheus, Grafana, and ELK or ...
$112K - $257K/yr
Experience with Infrastructure-as-Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Prometheus, Grafana, and ELK or ...
$112K - $257K/yr
Experience with Infrastructure-as-Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Prometheus, Grafana, and ELK or ...
$112K - $257K/yr
Experience with Infrastructure-as-Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Pro met heus, Grafana, and ELK ...
$112K - $257K/yr
Experience with Infrastructure-as-Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Pro met heus, Grafana, and ELK ...
$112K - $257K/yr
Experience with Infrastructure as Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Pro met heus, Grafana, and ELK ...
$112K - $257K/yr
Experience with Infrastructure as Code and configuration management using Terraform and Ansible * Experience implementing and operating an observability stack such as Pro met heus, Grafana, and ELK ...
$104K - $142K/yr
... observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments. We are ...
$104K - $142K/yr
... observability, capacity, and service enablement. We build software and automation to standardize how network platforms and services are deployed, scaled, and managed across environments. We are ...
$122K - $161K/yr
Cloud & DevOps * Deploy and manage applications within cloud environments such as: Azure, AWS ... Familiarity with observability tools, logging platforms, and performance monitoring. Education ...
$122K - $161K/yr
Cloud & DevOps * Deploy and manage applications within cloud environments such as: Azure, AWS ... Familiarity with observability tools, logging platforms, and performance monitoring. Education ...
... management within Kubernetes environments * Experience with containerization and orchestration technologies, monitoring, and observability solutions for AI deployments * Excellent knowledge of the ...
... management within Kubernetes environments * Experience with containerization and orchestration technologies, monitoring, and observability solutions for AI deployments * Excellent knowledge of the ...
$63.75 - $82/hr
Data Architect (S/4HANA) Sr. Manager, Enterprise Architecture & Strategy - Individual Contributor ... and observability. * Create and maintain architecture artifacts (capability maps, value streams ...
$63.75 - $82/hr
Data Architect (S/4HANA) Sr. Manager, Enterprise Architecture & Strategy - Individual Contributor ... and observability. * Create and maintain architecture artifacts (capability maps, value streams ...
$90K - $153K/yr
Optimize cloud resource utilization and cost management strategies. Qualifications: * 5 to 7 years ... Experience with monitoring & observability tools (CloudWatch, Datadog) EEO Policy Statement Pay ...
$90K - $153K/yr
Optimize cloud resource utilization and cost management strategies. Qualifications: * 5 to 7 years ... Experience with monitoring & observability tools (CloudWatch, Datadog) EEO Policy Statement Pay ...
$112K - $257K/yr
Experience implementing a secure SDLC, including scanning, secrets management, and secure coding ... Experience with observability for distributed services, such as Prometheus, Grafana, or ...
$112K - $257K/yr
Experience implementing a secure SDLC, including scanning, secrets management, and secure coding ... Experience with observability for distributed services, such as Prometheus, Grafana, or ...
| Aspect | Observability Manager | Site Reliability Engineer |
|---|---|---|
| Credentials | Typically requires experience in monitoring, logging, and cloud tools; certifications like AWS, Google Cloud, or Kubernetes are common | Requires strong background in systems engineering, scripting, and cloud platforms; certifications like AWS, GCP, or Linux are often preferred |
| Work Environment | Focuses on overseeing observability tools, data analysis, and team coordination in tech environments | Hands-on role involving system automation, incident response, and infrastructure reliability |
| Industry Usage | Used across tech companies to improve system visibility and performance | Common in DevOps and SRE teams to ensure system reliability and uptime |
The Observability Manager primarily oversees monitoring and logging strategies, ensuring system visibility, while the Site Reliability Engineer is more hands-on, focusing on automating infrastructure and maintaining system reliability. Both roles require technical expertise and often collaborate closely but differ in scope and daily responsibilities.

On-site, Remote
9.6
Based on 18 frontline employees who took The Breakroom Quiz
7th of 246 rated software companies
Great coworkers
People enjoy working here
Good employer
Respectful managers
Learn new skills
$108K - $147K/yr
Full-time
Posted 21 days ago
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology-and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent. As an NVIDIAN, you'll be immersed in a diverse, supportive environment where everyone is inspired to do their best work. Come join the team and see how you can make a lasting impact on the world.
NVIDIA is seeking an experienced software engineer to join the Cloud Foundations Automation team. Our team builds and operates the core infrastructure services that power NVIDIA's DGX Cloud and SuperPod deployments, delivering secure, reliable, and observable platforms at global scale.
What you'll be doing:
Build and operate core infrastructure services that power NVIDIA's global AI infrastructure.
Architect and develop secure, scalable and highly available cloud-native platform services.
Develop software that enables infrastructure orchestration, self-service workflows, and platform automation.
Own integrations with internal and external platforms to automate infrastructure provisioning and lifecycle management.
Build observability and security capabilities that improve the reliability and resilience of our infrastructure.
Partner with infrastructure and networking teams to deliver production services at scale.
Drive operational excellence through automation, monitoring, incident response, and continuous improvement.
What we need to see:
BS or equivalent experience with 8+ years of relevant industry experience.
Strong proficiency in Python and Go, with experience building production-quality software.
Experience building cloud-native microservices and APIs on Kubernetes using frameworks such as FastAPI, gRPC, or REST.
Experience with infrastructure automation (Terraform, Ansible), workflow orchestration (Temporal), and distributed systems using databases, Redis, and messaging platforms (Kafka, NATS, SQS).
Experience designing, building, and operating production infrastructure services such as DNS, NTP, AAA (RADIUS/OAuth), and observability platforms.
Strong Linux fundamentals with experience in observability (Prometheus, Grafana, OpenTelemetry, gNMI), networking (BGP, switching, routing, load balancing), and security (VPNs, firewalls, iptables/nftables).
Excellent problem-solving, communication, and collaboration skills.
Ways to stand out from the crowd:
Hands-on experience with network infrastructure including switches, routers, and firewalls.
Familiarity with InfiniBand, RDMA, and AI/HPC networking.Experience with NetBox, Nautobot, or similar network source of truth platforms.
Contributions to open-source software.Experience with public cloud platforms.
You will also be eligible for equity and benefits.
This posting is for an existing vacancy.
NVIDIA uses AI tools in its recruiting processes.
NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Computer and electronic product manufacturing
10,000+ Employees
Santa Clara, CA, US