NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI ... Preferred : • Technical competency in managing and automating large-scale distributed systems ...
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI ... Preferred : • Technical competency in managing and automating large-scale distributed systems ...
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI ... Technical competency in managing and automating large-scale distributed systems independent of ...
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI ... Technical competency in managing and automating large-scale distributed systems independent of ...
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI ... Technical competency in managing and automating large-scale distributed systems independent of ...
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI ... Technical competency in managing and automating large-scale distributed systems independent of ...
Power System Engineer II
Raleigh, NC · On-site
$83K - $101K/yr
SE Engineering, PC is a specialized power systems engineering firm trusted by industrial, healt ... Design and improve low‐ and medium‐voltage electrical distribution systems * Work with ...
Quick apply
Power System Engineer II
Raleigh, NC · On-site
$83K - $101K/yr
SE Engineering, PC is a specialized power systems engineering firm trusted by industrial, healt ... Design and improve low‐ and medium‐voltage electrical distribution systems * Work with ...
Principal Software Engineer - Distributed Systems / Cloud-Native SaaS
Raleigh, NC · Hybrid
$162K - $189K/yr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Principal Software Engineer - Distributed Systems / Cloud-Native SaaS
Raleigh, NC · Hybrid
$162K - $189K/yr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Quick apply
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Principal Software Engineer -- Distributed Systems / Cloud-Native SaaS
Raleigh, NC · On-site
$162 - $189/hr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Principal Software Engineer -- Distributed Systems / Cloud-Native SaaS
Raleigh, NC · On-site
$162 - $189/hr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Principal Software Engineer - Distributed Systems / Cloud-Native SaaS
Raleigh, NC · On-site +1
$162K - $189K/yr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Principal Software Engineer - Distributed Systems / Cloud-Native SaaS
Raleigh, NC · On-site +1
$162K - $189K/yr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Principal Software Engineer -- Distributed Systems / Cloud-Native SaaS
Raleigh, NC · Hybrid
$162K - $189K/yr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Principal Software Engineer -- Distributed Systems / Cloud-Native SaaS
Raleigh, NC · Hybrid
$162K - $189K/yr
Join us as a Principal Software Engineer -Distributed Systems / Cloud-Native SaaS and help us do what we do best: propelling business forward. This will be a hybrid role working between your home ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Software Engineer - Core Systems and Storage Roles (Multiple Individual Contributor Levels)
Morrisville, NC · On-site
$120K - $280K/yr
Engineers in these roles design, build, and optimize foundational components of NetApp's storage ... You will work on real-world problems involving filesystems, storage internals, distributed systems ...
Software Engineer - Core Systems and Storage Roles (Multiple Individual Contributor Levels)
Morrisville, NC · On-site
$120K - $280K/yr
Engineers in these roles design, build, and optimize foundational components of NetApp's storage ... You will work on real-world problems involving filesystems, storage internals, distributed systems ...
... Distribution Systems Engineer
... Distribution Systems Engineer
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Principal Systems Engineer
Durham, NC · On-site
Monitors Cloud computing, distributed applications, and databases. Participates in the development ... Bachelor's degree in Computer Science, Engineering, Information Technology, Information Systems, or ...
Principal Systems Engineer
Durham, NC · On-site
Monitors Cloud computing, distributed applications, and databases. Participates in the development ... Bachelor's degree in Computer Science, Engineering, Information Technology, Information Systems, or ...
Senior Software Engineer, Infrastructure Automation and Distributed Systems
Durham, NC · On-site
$104K - $142K/yr
We are seeking Systems Engineers and Software Engineers interested in building and running reliable ... Experience with infrastructure automation and distributed systems design developing tools for ...
Senior Software Engineer, Infrastructure Automation and Distributed Systems
Durham, NC · On-site
$104K - $142K/yr
We are seeking Systems Engineers and Software Engineers interested in building and running reliable ... Experience with infrastructure automation and distributed systems design developing tools for ...
System Administrator
Raleigh, NC · On-site
Deployments to distributed systems. Software packaging. Testing and working in engineering lab environments. Troubleshooting, identifying, analyzing, and solving problems at customer sites for POS ...
System Administrator
Raleigh, NC · On-site
Deployments to distributed systems. Software packaging. Testing and working in engineering lab environments. Troubleshooting, identifying, analyzing, and solving problems at customer sites for POS ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distinguished Technologist, Distributed Systems & Agentic Query Optimization This role has been ... Collaborate with product, customer engineering, and support teams to ensure reliability ...
Distributed Systems Engineer information
See Raleigh, NC salary details
$52K - $62K
2% of jobs
$62K - $72.1K
4% of jobs
$72.1K - $82.1K
7% of jobs
$82.1K - $92.1K
9% of jobs
$94.9K is the 25th percentile. Wages below this are outliers.
$92.1K - $102.2K
10% of jobs
$102.2K - $112.2K
7% of jobs
$112.2K - $122.2K
10% of jobs
The median wage is $123.9K / yr.
$122.2K - $132.2K
6% of jobs
$132.2K - $142.3K
3% of jobs
$152K is the 75th percentile. Wages above this are outliers.
$142.3K - $152.3K
17% of jobs
$152.3K - $162.3K
24% of jobs
$52K
$123.7K
$162.3K
How much do distributed systems engineer jobs pay per year?
What are the key skills and qualifications needed to thrive as a distributed systems engineer?
To thrive as a Distributed Systems Engineer, you need a strong background in computer science, experience with large-scale system design, and proficiency in languages such as Java, Go, or Python. Familiarity with cloud platforms (like AWS, GCP, or Azure), container orchestration tools (such as Kubernetes), and distributed databases is commonly required, and certifications in cloud computing can be advantageous. Strong problem-solving abilities, collaboration, and excellent communication skills help you navigate complex issues and work effectively across technical teams. These skills are fundamental for designing, implementing, and maintaining robust distributed systems that perform reliably at scale.
What does a distributed systems engineer do?
A Distributed Systems Engineer designs, builds, and maintains large-scale systems that run across multiple machines or data centers. They ensure reliability, scalability, and fault tolerance by using technologies like cloud computing, containerization, and distributed databases. Their work often involves solving complex problems related to data consistency, network latency, and system coordination.

Full-time
Re-posted 24 days ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
8th of 242 rated software companies
Job description
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. The role involves working on production systems for large scalable GPU clusters and implementing monitoring capabilities to ensure reliability and performance.
Responsibilities:
• You will be part of an DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI workloads. This includes working on custom software related to scheduling GPU resources on kubernetes.
• Implementing monitoring and health management capabilities that enable industry leading reliability, availability, and scalability of GPU assets. You will be harnessing multiple data streams, ranging from GPU hardware diagnostics to cluster and network telemetry.
• Working with teams across NVIDIA to ensure production AI clusters run reliability and consistently with maximum performance. Evaluating system failures and improving services based on a well-defined incident management process.
Qualifications:
Required:
• Significant software engineering experience with kubernetes including cluster operations, operator development, node health monitoring and working with GPU resource scheduling.
• Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work.
• Software development experience with kubernetes APIs and frameworks not just operating a cluster.
• Highly motivated with strong communication skills, you can work successfully with multi-functional teams, principles, and architects and coordinate effectively across organizational boundaries and geographies.
• 15+ years in similar role and experience on large-scale production systems.
• Experience with common software engineering principles, tools and techniques.
• You possess a BS in Computer Science, Engineering, Physics, Mathematics or a comparable Degree or equivalent experience.
• Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms.
Preferred:
• Technical competency in managing and automating large-scale distributed systems independent of cloud providers.
• Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Bright Cluster Manager).
• Proven operational excellence in maintaining reliable and performant AI infrastructure.
Company:
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.
About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993