1

Manager Supercomputer Jobs in Seattle, WA (NOW HIRING)

Senior GPU Supercomputer Scheduler Engineer

Redmond, WA · On-site

$137K - $180K/yr

... management and orchestration services • Provide support to staff and end users to resolve batch scheduler issues • Build and improve our ecosystem around GPU-accelerated computing • Performance ...

Software Development Manager , EC2 Nitro

Seattle, WA · On-site

$140K - $185K/yr

You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...

Software Development Manager , EC2 Nitro

Seattle, WA · On-site

$140K - $185K/yr

You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...

... centers, supercomputers, and desktop computers. We design and manufacture solutions used by the ... Lead review and negotiation of customer contracts and manage inputs from stakeholders to drive ...

Manager, Key Accounts ABOUT COOLIT SYSTEMS INC. Founded in Calgary, Alberta in 2001, CoolIT Systems ... centers, supercomputers, and desktop computers. We design and manufacture solutions used by the ...

Software Development Manager , EC2 Nitro

Seattle, WA · On-site

$140K - $185K/yr

You'll build and manage a team focused on establishing EC2 as the definitive source for ML ... Working with us means having the opportunity to influence the future of supercomputing in the cloud ...

next page

Showing results 1-20

Manager Supercomputer information

What is the difference between Manager Supercomputer vs Supercomputing Systems Engineer?

AspectManager SupercomputerSupercomputing Systems Engineer
Required CredentialsBachelor's or master's in computer science, engineering, or related field; management experienceBachelor's or master's in computer science, computer engineering, or related field; technical certifications
Work EnvironmentOversees supercomputing facilities, manages teams, strategic planningDesigns, develops, and maintains supercomputing systems, works hands-on with hardware/software
Employer & Industry UsageResearch labs, government agencies, large tech companiesResearch institutions, high-performance computing centers, tech firms

The Manager Supercomputer primarily oversees supercomputing operations and manages teams, focusing on strategic and administrative tasks. In contrast, the Supercomputing Systems Engineer is more technically involved, designing and maintaining supercomputing systems. Both roles require strong technical backgrounds, but their responsibilities differ in scope and focus.

What are popular job titles related to Manager Supercomputer jobs in Seattle, WA? For Manager Supercomputer jobs in Seattle, WA, the most frequently searched job titles are:
What job categories do people searching Manager Supercomputer jobs in Seattle, WA look for? The top searched job categories for Manager Supercomputer jobs in Seattle, WA are:
Infographic showing various Manager Supercomputer job openings in Seattle, WA as of July 2026, with employment types broken down into 81% Full Time, 16% Part Time, 1% Temporary, and 2% Contract. Highlights an 86% Physical, 1% Hybrid, and 13% Remote job distribution.

Member of Technical Staff - Compute Infrastructure

xAI

Seattle, WA • On-site

Full-time

Re-posted 28 days ago


Job description

Job Summary:
xAI is dedicated to creating AI systems that enhance humanity's understanding of the universe. The role involves owning and optimizing GPU supercomputers and the platform layer for training and inference, requiring a blend of low-level systems programming and high-scale infrastructure work.
Responsibilities:
• Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads
• Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance
• Work on Linux kernel internals, scheduling, memory management, and resource isolation at cluster scale
• Build custom container orchestration, virtualization layers (KVM, Firecracker, etc.), and distributed systems that go beyond standard Kubernetes
• Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations
• Create and maintain infrastructure-as-code, automation, and tools that keep the entire supercomputer reliable and efficient
• Collaborate closely with AI research teams to deliver production-grade performance and scalability
Qualifications:
Required:
• Design, build, and optimize massive GPU clusters for extreme-scale training and inference workloads
• Develop and tune low-level CUDA kernels (GeMM, Attention, etc.), using CUTLASS, Tensor Cores, and Nsight for maximum performance
• Work on Linux kernel internals, scheduling, memory management, and resource isolation at cluster scale
• Build custom container orchestration, virtualization layers (KVM, Firecracker, etc.), and distributed systems that go beyond standard Kubernetes
• Profile, debug, and eliminate bottlenecks across GPU memory hierarchy, networking fabric, filesystems, and multi-GPU operations
• Create and maintain infrastructure-as-code, automation, and tools that keep the entire supercomputer reliable and efficient
• Collaborate closely with AI research teams to deliver production-grade performance and scalability
Preferred:
• Deep low-level systems programming (C/C++ or Rust)
• Experience building and operating high performance exabyte scale storage systems
• Strong experience with large-scale GPU clusters or distributed compute infrastructure at production scale
• Hands-on work with GPU kernel optimization (CUTLASS, custom kernels, Nsight profiling)
• Experience with Linux kernel internals, scheduling, virtualization, or large-scale orchestration
• Track record of building or running high-performance infrastructure for AI workloads (training or inference platforms)
• Ability to reason from first principles and optimize for both memory-bound and compute-bound scenarios
Company:
Understand the Universe. Founded in 2023, the company is headquartered in , , with a team of 1001-5000 employees. The company is currently Late Stage.