Preferred : • Technical competency in managing and automating large-scale distributed systems independent of cloud providers. • Advanced hands-on experience and deep understanding of cluster ...
Preferred : • Technical competency in managing and automating large-scale distributed systems independent of cloud providers. • Advanced hands-on experience and deep understanding of cluster ...
Member of Technical Staff - Cluster Infrastructure & Supercomputing
Palo Alto, CA · On-site
$200K - $400K/yr
Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers) * Hands-on experience with GPU/TPU infrastructure in production environments * Strong Linux systems and ...
Member of Technical Staff - Cluster Infrastructure & Supercomputing
Palo Alto, CA · On-site
$200K - $400K/yr
Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers) * Hands-on experience with GPU/TPU infrastructure in production environments * Strong Linux systems and ...
The Critical Projects Implementation (CPI) team is a project management and execution team that manages construction activity within the operational data center spaces. The CPI team is tasked with ...
New
The Critical Projects Implementation (CPI) team is a project management and execution team that manages construction activity within the operational data center spaces. The CPI team is tasked with ...
New
The Critical Projects Implementation (CPI) team is a project management and execution team that manages construction activity within the operational data center spaces. The CPI team is tasked with ...
New
The Critical Projects Implementation (CPI) team is a project management and execution team that manages construction activity within the operational data center spaces. The CPI team is tasked with ...
New
Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers) * Hands-on experience with GPU/TPU infrastructure in production environments * Strong Linux systems and ...
Deep experience with cluster management systems (Kubernetes, Slurm, Ray, or custom schedulers) * Hands-on experience with GPU/TPU infrastructure in production environments * Strong Linux systems and ...
Own the cluster management user interface, including provisioning flows, natural language interfaces, and real-time proactive system insights. * Design and build visualization systems that make ...
Own the cluster management user interface, including provisioning flows, natural language interfaces, and real-time proactive system insights. * Design and build visualization systems that make ...
Member of Technical Staff, Infrastructure
San Francisco, CA · On-site
$200K - $280K/yr
Medical
Dental
Vision
You'll write Go for control-plane services like cluster-manager, traffic-control-plane, and environment-manager, and you'll set the bar for how Vapi runs stateful workloads at scale. What You'll Do ...
Member of Technical Staff, Infrastructure
San Francisco, CA · On-site
$200K - $280K/yr
Medical
Dental
Vision
You'll write Go for control-plane services like cluster-manager, traffic-control-plane, and environment-manager, and you'll set the bar for how Vapi runs stateful workloads at scale. What You'll Do ...
System Software Engineer - Node & Cluster Management
Mountain View, CA · On-site
$175K - $362K/yr
Medical
Dental
Vision
Life
Retirement
PTO
Design and implement cluster management solutions and failover algorithms to minimize downtime * Build the management CLI utilities that operators and internal engineers use daily - interacting with ...
System Software Engineer - Node & Cluster Management
Mountain View, CA · On-site
$175K - $362K/yr
Medical
Dental
Vision
Life
Retirement
PTO
Design and implement cluster management solutions and failover algorithms to minimize downtime * Build the management CLI utilities that operators and internal engineers use daily - interacting with ...
Own the cluster management user interface, including provisioning flows, natural language interfaces, and real-time proactive system insights. * Design and build visualization systems that make ...
Quick apply
Own the cluster management user interface, including provisioning flows, natural language interfaces, and real-time proactive system insights. * Design and build visualization systems that make ...
System Software Engineer - Node & Cluster Management
$175K - $362K/yr
Medical
Dental
Vision
Life
Retirement
PTO
Design and implement cluster management solutions and failover algorithms to minimize downtime * Build the management CLI utilities that operators and internal engineers use daily - interacting with ...
System Software Engineer - Node & Cluster Management
$175K - $362K/yr
Medical
Dental
Vision
Life
Retirement
PTO
Design and implement cluster management solutions and failover algorithms to minimize downtime * Build the management CLI utilities that operators and internal engineers use daily - interacting with ...
Cluster Telecom Service Manager (CTSM)
Los Angeles, CA · On-site
Medical
Dental
Vision
Life
Retirement
PTO
The Cluster Telecommunications Service Manager (CTSM) is responsible for the end-to-end strategic planning, coordination, and intentional delivery of telecommunications services across a defined ...
Cluster Telecom Service Manager (CTSM)
Los Angeles, CA · On-site
Medical
Dental
Vision
Life
Retirement
PTO
The Cluster Telecommunications Service Manager (CTSM) is responsible for the end-to-end strategic planning, coordination, and intentional delivery of telecommunications services across a defined ...
Cluster Telecom Service Manager (CTSM)
$70 - $80/hr
Medical
Dental
Vision
Life
Retirement
PTO
The Cluster Telecommunications Service Manager (CTSM) is responsible for the end-to-end strategic planning, coordination, and intentional delivery of telecommunications services across a defined ...
Cluster Telecom Service Manager (CTSM)
$70 - $80/hr
Medical
Dental
Vision
Life
Retirement
PTO
The Cluster Telecommunications Service Manager (CTSM) is responsible for the end-to-end strategic planning, coordination, and intentional delivery of telecommunications services across a defined ...
Cluster Telecom Service Manager (CTSM)
Medical
Dental
Vision
Life
Retirement
PTO
The Cluster Telecommunications Service Manager (CTSM) is responsible for the end-to-end strategic planning, coordination, and intentional delivery of telecommunications services across a defined ...
Cluster Telecom Service Manager (CTSM)
Medical
Dental
Vision
Life
Retirement
PTO
The Cluster Telecommunications Service Manager (CTSM) is responsible for the end-to-end strategic planning, coordination, and intentional delivery of telecommunications services across a defined ...
Sr. Splunk engineer
Croydon, PA · On-site
$67/hr
Design and implement a multi-site, highly available Splunk Enterprise deployment including Cluster Manager, License Master, Deployer, * Deployment Server, Monitoring Console, multi-site indexer ...
Posted today
Sr. Splunk engineer
Croydon, PA · On-site
$67/hr
Design and implement a multi-site, highly available Splunk Enterprise deployment including Cluster Manager, License Master, Deployer, * Deployment Server, Monitoring Console, multi-site indexer ...
Posted today
Sr. Splunk engineer
Croydon, PA · On-site
$67/hr
Design and implement a multi-site, highly available Splunk Enterprise deployment including Cluster Manager, License Master, Deployer, * Deployment Server, Monitoring Console, multi-site indexer ...
Posted today
Sr. Splunk engineer
Croydon, PA · On-site
$67/hr
Design and implement a multi-site, highly available Splunk Enterprise deployment including Cluster Manager, License Master, Deployer, * Deployment Server, Monitoring Console, multi-site indexer ...
Posted today
Expert System Administrator (Virtual Infrastructure & Cluster Management)
College Park, MD · On-site +1
Medical
Dental
Vision
Life
Retirement
PTO
Expert System Administrator (Virtual Infrastructure & Cluster Management) Location: College Park, MD Compensation Range: $220 - $240K *Clearance: *Active TS/SCI w/ Polygraph needed to apply * Company ...
Expert System Administrator (Virtual Infrastructure & Cluster Management)
College Park, MD · On-site +1
Medical
Dental
Vision
Life
Retirement
PTO
Expert System Administrator (Virtual Infrastructure & Cluster Management) Location: College Park, MD Compensation Range: $220 - $240K *Clearance: *Active TS/SCI w/ Polygraph needed to apply * Company ...
Senior AI Cluster Hardware Engineer
Austin, TX · On-site
$142K/yr
Familiarity with cluster management tools and systems. * Excellent communication and collaboration ... skills for effective teamwork. * RDMA network configuration, troubleshooting and performance tuning.
Senior AI Cluster Hardware Engineer
Austin, TX · On-site
$142K/yr
Familiarity with cluster management tools and systems. * Excellent communication and collaboration ... skills for effective teamwork. * RDMA network configuration, troubleshooting and performance tuning.
Data Center Operations Cluster Manager, DCC Communities
New Carlisle, IN · On-site
$152K/yr
This leader will direct managers responsible for mission-critical data center infrastructure, manage network systems, server operations, and infrastructure maintenance, and ensure optimal ...
Posted today
Data Center Operations Cluster Manager, DCC Communities
New Carlisle, IN · On-site
$152K/yr
This leader will direct managers responsible for mission-critical data center infrastructure, manage network systems, server operations, and infrastructure maintenance, and ensure optimal ...
Posted today
Cluster Engineer
Champaign, IL · On-site
Application Development Project Management Quality Assurance Business/Systems Analysis ... Cluster Engineer Job Details The main duties are building, configuring, and maintaining high ...
Cluster Engineer
Champaign, IL · On-site
Application Development Project Management Quality Assurance Business/Systems Analysis ... Cluster Engineer Job Details The main duties are building, configuring, and maintaining high ...
Senior AI Cluster Hardware Engineer
Austin, TX · Hybrid
$109K - $146K/yr
Familiarity with cluster management tools and systems. * Excellent communication and collaboration skills for effective teamwork. * RDMA network configuration, troubleshooting and performance tuning.
Senior AI Cluster Hardware Engineer
Austin, TX · Hybrid
$109K - $146K/yr
Familiarity with cluster management tools and systems. * Excellent communication and collaboration skills for effective teamwork. * RDMA network configuration, troubleshooting and performance tuning.
Cluster Manager information
See salary details
$29K - $37.1K
1% of jobs
$37.1K - $45.2K
5% of jobs
$45.2K - $53.3K
3% of jobs
$53.3K - $61.4K
3% of jobs
$61.4K - $69.5K
3% of jobs
$69.5K - $77.5K
1% of jobs
$77.5K - $85.6K
1% of jobs
$85.6K - $93.7K
1% of jobs
$93.7K - $101.8K
1% of jobs
$101.8K - $109.9K
1% of jobs
$110.3K is the 25th percentile. Wages below this are outliers.
$109.9K - $118K
79% of jobs
$29K
$104.6K
$118K
How much do cluster manager jobs pay per year?
What is the difference between Cluster Manager vs Operations Manager?
| Aspect | Cluster Manager | Operations Manager |
|---|---|---|
| Credentials | Bachelor's degree in business, management, or related field; certifications like PMP are common | Bachelor's degree in business, management, or related field; certifications like PMP are common |
| Work Environment | Oversees multiple locations or units within a region or sector | Manages daily operations within a specific department or facility |
| Industry Usage | Common in retail, healthcare, logistics, and hospitality sectors | Widely used across various industries including manufacturing, retail, and services |
While both roles involve management responsibilities, a Cluster Manager oversees multiple sites or units within a region, focusing on strategic coordination, whereas an Operations Manager concentrates on the daily functioning of a specific department or location. The roles often overlap but differ mainly in scope and scale of oversight.
What is the cluster manager?
What do you need to be a cluster manager?

Full-time
Re-posted 15 hours ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
7th of 243 rated software companies
Job description
NVIDIA is hiring experienced software engineers with kubernetes experience to help scale up its AI Infrastructure. The role involves working on production systems for large scalable GPU clusters and implementing monitoring capabilities to ensure reliability and performance.
Responsibilities:
• You will be part of an DGX Cloud team responsible for production systems that enable large scalable GPU clusters to be used for a variety of AI workloads. This includes working on custom software related to scheduling GPU resources on kubernetes.
• Implementing monitoring and health management capabilities that enable industry leading reliability, availability, and scalability of GPU assets. You will be harnessing multiple data streams, ranging from GPU hardware diagnostics to cluster and network telemetry.
• Working with teams across NVIDIA to ensure production AI clusters run reliability and consistently with maximum performance. Evaluating system failures and improving services based on a well-defined incident management process.
Qualifications:
Required:
• Significant software engineering experience with kubernetes including cluster operations, operator development, node health monitoring and working with GPU resource scheduling.
• Direct experience in a software engineering role within a highly technical organization with demonstrable impact from your work.
• Software development experience with kubernetes APIs and frameworks not just operating a cluster.
• Highly motivated with strong communication skills, you can work successfully with multi-functional teams, principles, and architects and coordinate effectively across organizational boundaries and geographies.
• 15+ years in similar role and experience on large-scale production systems.
• Experience with common software engineering principles, tools and techniques.
• You possess a BS in Computer Science, Engineering, Physics, Mathematics or a comparable Degree or equivalent experience.
• Technical knowledge, including a systems programming language (Go, Python) and a solid understanding of data structures and algorithms.
Preferred:
• Technical competency in managing and automating large-scale distributed systems independent of cloud providers.
• Advanced hands-on experience and deep understanding of cluster management systems (Kubernetes, Slurm, Bright Cluster Manager).
• Proven operational excellence in maintaining reliable and performant AI infrastructure.
Company:
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.
About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993