The role involves owning and optimizing GPU supercomputers and the platform layer for training and ... management, and resource isolation at cluster scale • Build custom container orchestration ...
The role involves owning and optimizing GPU supercomputers and the platform layer for training and ... management, and resource isolation at cluster scale • Build custom container orchestration ...
The role involves building and optimizing GPU supercomputers and the platform layer for AI training ... management, and resource isolation at cluster scale • Build custom container orchestration ...
The role involves building and optimizing GPU supercomputers and the platform layer for AI training ... management, and resource isolation at cluster scale • Build custom container orchestration ...
Senior GPU Supercomputer Scheduler Engineer
$137K - $180K/yr
Within this mission, our team, Managed AI Research Superclusters (MARS), builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next ...
Senior GPU Supercomputer Scheduler Engineer
$137K - $180K/yr
Within this mission, our team, Managed AI Research Superclusters (MARS), builds and scales the infrastructure, platforms, and tools that enable researchers and engineers to develop the next ...
Senior GPU Supercomputer Scheduler Engineer
Redmond, WA · On-site
$137K - $180K/yr
... management and orchestration services • Provide support to staff and end users to resolve batch scheduler issues • Build and improve our ecosystem around GPU-accelerated computing • Performance ...
Senior GPU Supercomputer Scheduler Engineer
Redmond, WA · On-site
$137K - $180K/yr
... management and orchestration services • Provide support to staff and end users to resolve batch scheduler issues • Build and improve our ecosystem around GPU-accelerated computing • Performance ...
Software Development Manager , EC2 Nitro
Seattle, WA · On-site
$140K - $185K/yr
You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...
Software Development Manager , EC2 Nitro
Seattle, WA · On-site
$140K - $185K/yr
You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...
Software Development Manager , EC2 Nitro
Seattle, WA · On-site
$140K - $185K/yr
You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...
Software Development Manager , EC2 Nitro
Seattle, WA · On-site
$140K - $185K/yr
You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...
... management (Python, Bash, Ansible, or similar) * Experience designing AI supercomputers from MEP designs * Experience delivering bare-metal infrastructure at scale for large strategic customers or AI ...
Quick apply
... management (Python, Bash, Ansible, or similar) * Experience designing AI supercomputers from MEP designs * Experience delivering bare-metal infrastructure at scale for large strategic customers or AI ...
Senior Specialist Field Engineer - Compute Infrastructure
Bellevue, WA · On-site
$122K - $166K/yr
... management (Python, Bash, Ansible, or similar) * Experience designing AI supercomputers from MEP designs * Experience delivering bare-metal infrastructure at scale for large strategic customers or AI ...
Senior Specialist Field Engineer - Compute Infrastructure
Bellevue, WA · On-site
$122K - $166K/yr
... management (Python, Bash, Ansible, or similar) * Experience designing AI supercomputers from MEP designs * Experience delivering bare-metal infrastructure at scale for large strategic customers or AI ...
Operations Engineer, Fleet Reliability
Bellevue, WA · On-site +1
$83K - $110K/yr
Configure and maintain large-scale high-performance supercomputing clusters running ... practices for system management * Think critically about your day-to-day work and work ...
Quick apply
Operations Engineer, Fleet Reliability
Bellevue, WA · On-site +1
$83K - $110K/yr
Configure and maintain large-scale high-performance supercomputing clusters running ... practices for system management * Think critically about your day-to-day work and work ...
Key Account Manager
Seattle, WA · On-site
... centers, supercomputers, and desktop computers. We design and manufacture solutions used by the ... Lead review and negotiation of customer contracts and manage inputs from stakeholders to drive ...
Key Account Manager
Seattle, WA · On-site
... centers, supercomputers, and desktop computers. We design and manufacture solutions used by the ... Lead review and negotiation of customer contracts and manage inputs from stakeholders to drive ...
Communications Specialist, Chips, Amazon
$60K - $80K/yr
... AI supercomputers. As PR Specialist for Chips, you'll support how customers, media, and the ... You'll need strong writing skills, good project management instincts, and the initiative to see ...
Communications Specialist, Chips, Amazon
$60K - $80K/yr
... AI supercomputers. As PR Specialist for Chips, you'll support how customers, media, and the ... You'll need strong writing skills, good project management instincts, and the initiative to see ...
Manager, Key Accounts US
Seattle, WA · On-site
Manager, Key Accounts ABOUT COOLIT SYSTEMS INC. Founded in Calgary, Alberta in 2001, CoolIT Systems ... centers, supercomputers, and desktop computers. We design and manufacture solutions used by the ...
Manager, Key Accounts US
Seattle, WA · On-site
Manager, Key Accounts ABOUT COOLIT SYSTEMS INC. Founded in Calgary, Alberta in 2001, CoolIT Systems ... centers, supercomputers, and desktop computers. We design and manufacture solutions used by the ...
Communications Specialist, Chips, Amazon
$60K - $80K/yr
... AI supercomputers. As PR Specialist for Chips, you'll support how customers, media, and the ... You'll need strong writing skills, good project management instincts, and the initiative to see ...
Communications Specialist, Chips, Amazon
$60K - $80K/yr
... AI supercomputers. As PR Specialist for Chips, you'll support how customers, media, and the ... You'll need strong writing skills, good project management instincts, and the initiative to see ...
Operations Engineer, Fleet Reliability
Bellevue, WA · On-site
$83K - $110K/yr
Configure and maintain large-scale high-performance supercomputing clusters running ... practices for system management * Think critically about your day-to-day work and work ...
Operations Engineer, Fleet Reliability
Bellevue, WA · On-site
$83K - $110K/yr
Configure and maintain large-scale high-performance supercomputing clusters running ... practices for system management * Think critically about your day-to-day work and work ...
Software Development Manager , EC2 Nitro
Seattle, WA · On-site
$140K - $185K/yr
You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...
Software Development Manager , EC2 Nitro
Seattle, WA · On-site
$140K - $185K/yr
You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
Senior Software Dev Engineer, EC2 Nitro, EC2 Nitro
Seattle, WA · On-site
$139K - $183K/yr
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
Senior Software Dev Engineer, EC2 Nitro, EC2 Nitro
Seattle, WA · On-site
$139K - $183K/yr
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
Senior Software Dev Engineer, EC2 Nitro, EC2 Nitro
$139K - $183K/yr
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
Senior Software Dev Engineer, EC2 Nitro, EC2 Nitro
$139K - $183K/yr
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
... of EC2 Supercomputers, optimized for high-performance training and inference workloads. We are ... Experience with Linux package management, version control systems, automated build processes, and ...
Manager Supercomputer information
What is the difference between Manager Supercomputer vs Supercomputing Systems Engineer?
| Aspect | Manager Supercomputer | Supercomputing Systems Engineer |
|---|---|---|
| Required Credentials | Bachelor's or master's in computer science, engineering, or related field; management experience | Bachelor's or master's in computer science, computer engineering, or related field; technical certifications |
| Work Environment | Oversees supercomputing facilities, manages teams, strategic planning | Designs, develops, and maintains supercomputing systems, works hands-on with hardware/software |
| Employer & Industry Usage | Research labs, government agencies, large tech companies | Research institutions, high-performance computing centers, tech firms |
The Manager Supercomputer primarily oversees supercomputing operations and manages teams, focusing on strategic and administrative tasks. In contrast, the Supercomputing Systems Engineer is more technically involved, designing and maintaining supercomputing systems. Both roles require strong technical backgrounds, but their responsibilities differ in scope and focus.

Full-time
Re-posted 28 days ago
Job description
xAI is dedicated to creating AI systems that enhance humanity's understanding of the universe. The role involves owning and optimizing GPU supercomputers and the platform layer for training and inference, requiring a blend of low-level systems programming and high-scale infrastructure work.
Responsibilities:
• Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads
• Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance
• Work on Linux kernel internals, scheduling, memory management, and resource isolation at cluster scale
• Build custom container orchestration, virtualization layers (KVM, Firecracker, etc.), and distributed systems that go beyond standard Kubernetes
• Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations
• Create and maintain infrastructure-as-code, automation, and tools that keep the entire supercomputer reliable and efficient
• Collaborate closely with AI research teams to deliver production-grade performance and scalability
Qualifications:
Required:
• Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads
• Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance
• Work on Linux kernel internals, scheduling, memory management, and resource isolation at cluster scale
• Build custom container orchestration, virtualization layers (KVM, Firecracker, etc.), and distributed systems that go beyond standard Kubernetes
• Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations
• Create and maintain infrastructure-as-code, automation, and tools that keep the entire supercomputer reliable and efficient
• Collaborate closely with AI research teams to deliver production-grade performance and scalability
Preferred:
• Deep low-level systems programming (C/C++ or Rust)
• Experience building and operating high performance exabyte scale storage systems
• Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale
• Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling)
• Experience with Linux kernel internals, scheduling, virtualization, or large-scale orchestration
• Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms)
• Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios
Company:
Understand the Universe. Founded in 2023, the company is headquartered in , , with a team of 1001-5000 employees. The company is currently Late Stage.