The Software Engineer, Compute (GPU) will own the health of the compute fleet, build automation for deployment and repair, and ensure the reliability and scalability of the GPU fleet.
The Software Engineer, Compute (GPU) will own the health of the compute fleet, build automation for deployment and repair, and ensure the reliability and scalability of the GPU fleet.
Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
Sunnyvale, CA · On-site
$129K - $173K/yr
Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or ...
Quick apply
Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
Sunnyvale, CA · On-site
$129K - $173K/yr
Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or ...
Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
San Francisco, CA · On-site
$130K - $173K/yr
Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or ...
Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
San Francisco, CA · On-site
$130K - $173K/yr
Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or ...
Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
San Francisco, CA · On-site
$130K - $173K/yr
Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or ...
Quick apply
Senior Staff Data Center Operations Engineer, GPU Hardware Architecture
San Francisco, CA · On-site
$130K - $173K/yr
Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or ...
Principal Software Engineer, GPU Compute
$153K - $206K/yr
Drive GPU reliability and performance at fleet scale, defining the detection, diagnosis, and automated repair of unhealthy accelerators before they impact production. * Evaluate and onboard new GPU ...
Principal Software Engineer, GPU Compute
$153K - $206K/yr
Drive GPU reliability and performance at fleet scale, defining the detection, diagnosis, and automated repair of unhealthy accelerators before they impact production. * Evaluate and onboard new GPU ...
Software Engineer, Compute (GPU)
San Francisco, CA · On-site
$175K - $300K/yr
Build the repair pipeline that keeps pace with a fleet of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput problem. We're building the automation that takes a chip ...
Software Engineer, Compute (GPU)
San Francisco, CA · On-site
$175K - $300K/yr
Build the repair pipeline that keeps pace with a fleet of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput problem. We're building the automation that takes a chip ...
Software Engineer, GPU Infrastructure
San Francisco, CA · On-site
$175K - $300K/yr
Build the repair pipeline that keeps pace with a fleet of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput problem. We're building the automation that takes a chip ...
Software Engineer, GPU Infrastructure
San Francisco, CA · On-site
$175K - $300K/yr
Build the repair pipeline that keeps pace with a fleet of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput problem. We're building the automation that takes a chip ...
Principal Software Engineer, GPU Compute
San Mateo, CA · On-site
$345K - $399K/yr
Drive GPU reliability and performance at fleet scale, defining the detection, diagnosis, and automated repair of unhealthy accelerators before they impact production. * Evaluate and onboard new GPU ...
Principal Software Engineer, GPU Compute
San Mateo, CA · On-site
$345K - $399K/yr
Drive GPU reliability and performance at fleet scale, defining the detection, diagnosis, and automated repair of unhealthy accelerators before they impact production. * Evaluate and onboard new GPU ...
Senior HPC & GPU Infrastructure Engineer
San Francisco, CA · On-site
$150K - $220K/yr
Coordinate with data center staff, hardware vendors, and on-site technicians for repairs, RMA ... Lead deployment of new GPU nodes, including BIOS configuration, NUMA tuning, GPU topology ...
Senior HPC & GPU Infrastructure Engineer
San Francisco, CA · On-site
$150K - $220K/yr
Coordinate with data center staff, hardware vendors, and on-site technicians for repairs, RMA ... Lead deployment of new GPU nodes, including BIOS configuration, NUMA tuning, GPU topology ...
... GPU platforms. This will involve taking data from system logs, kernel logs, BMC redfish APIs, and ... Automate routine processes and build hardware diagnostics, provisioning and repair tooling Build ...
... GPU platforms. This will involve taking data from system logs, kernel logs, BMC redfish APIs, and ... Automate routine processes and build hardware diagnostics, provisioning and repair tooling Build ...
Infra Repair Technician
Springfield, NE · On-site
Infra Repair Technician Location - Springfield, NE 68059 Job type: Contract Note : May need to work ... Server hardware (e.g., GPU, CPU, Motherboard) * Storage systems (e.g., Hard Drives, SSDs) * Network ...
Quick apply
Infra Repair Technician
Springfield, NE · On-site
Infra Repair Technician Location - Springfield, NE 68059 Job type: Contract Note : May need to work ... Server hardware (e.g., GPU, CPU, Motherboard) * Storage systems (e.g., Hard Drives, SSDs) * Network ...
Debug Repair Technicians
Memphis, TN · On-site
$25 - $28/hr
As an Debug Repair Technician you will: * Perform advanced board-level troubleshooting and ... GPU, FPGA, ASIC, memory, and power subsystems Diagnose high speed digital, power distribution, and ...
Debug Repair Technicians
Memphis, TN · On-site
$25 - $28/hr
As an Debug Repair Technician you will: * Perform advanced board-level troubleshooting and ... GPU, FPGA, ASIC, memory, and power subsystems Diagnose high speed digital, power distribution, and ...
Staff Software Engineer, GPU Infrastructure Lifecycle Management
San Francisco, CA · On-site
$203K - $241K/yr
... up to GPU driver/CUDA stack, health validation, and decommission/RMA - as explicit, versioned ... Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or ...
Staff Software Engineer, GPU Infrastructure Lifecycle Management
San Francisco, CA · On-site
$203K - $241K/yr
... up to GPU driver/CUDA stack, health validation, and decommission/RMA - as explicit, versioned ... Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or ...
Staff Software Engineer, GPU Infrastructure Lifecycle Management
San Francisco, CA · On-site
$203K - $241K/yr
... up to GPU driver/CUDA stack, health validation, and decommission/RMA - as explicit, versioned ... Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or ...
Staff Software Engineer, GPU Infrastructure Lifecycle Management
San Francisco, CA · On-site
$203K - $241K/yr
... up to GPU driver/CUDA stack, health validation, and decommission/RMA - as explicit, versioned ... Automate self-healing: detect degraded or failed nodes, drain them safely, trigger repair or ...
Facilities Maintenance (GSE) Technician II
Kansas City, MO · On-site
$17.25 - $23.75/hr
Aircraft 400Hz Generator (GPU) repairs * Solid State GPU's. Transformers. * Hot work/welding and solder * Love efficiency and can help build improvements in processes * Dream about airplanes Other ...
Quick apply
Facilities Maintenance (GSE) Technician II
Kansas City, MO · On-site
$17.25 - $23.75/hr
Aircraft 400Hz Generator (GPU) repairs * Solid State GPU's. Transformers. * Hot work/welding and solder * Love efficiency and can help build improvements in processes * Dream about airplanes Other ...
Infrastructure Repair Technician
Huntsville, AL · On-site
$25 - $30/hr
Infra Repair Technician Location - Huntsville, AL, US 35810 Note - May need to work in shifts ... Server hardware (e.g., GPU, CPU, Motherboard) * Storage systems (e.g., Hard Drives, SSDs) * Network ...
Quick apply
Infrastructure Repair Technician
Huntsville, AL · On-site
$25 - $30/hr
Infra Repair Technician Location - Huntsville, AL, US 35810 Note - May need to work in shifts ... Server hardware (e.g., GPU, CPU, Motherboard) * Storage systems (e.g., Hard Drives, SSDs) * Network ...
Site Reliability Engineer, Compute
San Francisco, CA · On-site
$67.25 - $89.25/hr
The Site Reliability Engineer will be responsible for ensuring the health and reliability of the compute fleet, developing automation for deployment and repair processes, and qualifying new GPU ...
Site Reliability Engineer, Compute
San Francisco, CA · On-site
$67.25 - $89.25/hr
The Site Reliability Engineer will be responsible for ensuring the health and reliability of the compute fleet, developing automation for deployment and repair processes, and qualifying new GPU ...
Associate Product Manager
San Jose, CA · On-site
$90K - $105K/yr
The Associate Product Manager will support and coordinate GPU-related operations across RMA, Service, Repair, Technical Support, and Production teams, with a focus on improving workflow efficiency ...
Associate Product Manager
San Jose, CA · On-site
$90K - $105K/yr
The Associate Product Manager will support and coordinate GPU-related operations across RMA, Service, Repair, Technical Support, and Production teams, with a focus on improving workflow efficiency ...
Perform repairs and replacements of faulty components, including but not limited to: • Server hardware (e.g., GPU, CPU, Motherboard) • Storage systems (e.g., Hard Drives, SSDs) • Network ...
Quick apply
Perform repairs and replacements of faulty components, including but not limited to: • Server hardware (e.g., GPU, CPU, Motherboard) • Storage systems (e.g., Hard Drives, SSDs) • Network ...
Infrastructure Repair Technician - Datacenter
Huntsville, AL · On-site
$28/hr
Perform repairs and replacements of faulty components, including but not limited to server hardware (e.g., GPU, CPU, Motherboard), storage systems (e.g., Hard Drives, SSDs), and network devices (e.g ...
Quick apply
Infrastructure Repair Technician - Datacenter
Huntsville, AL · On-site
$28/hr
Perform repairs and replacements of faulty components, including but not limited to server hardware (e.g., GPU, CPU, Motherboard), storage systems (e.g., Hard Drives, SSDs), and network devices (e.g ...
Gpu Repair information
See salary details
$12.50 - $14.29
2% of jobs
$14.29 - $16.08
7% of jobs
$17.82 is the 25th percentile. Wages below this are outliers.
$16.08 - $17.88
16% of jobs
$17.88 - $19.67
21% of jobs
The median wage is $20 / hr.
$19.67 - $21.46
17% of jobs
$21.46 - $23.25
11% of jobs
$23.34 is the 75th percentile. Wages above this are outliers.
$23.25 - $25.04
11% of jobs
$25.04 - $26.84
6% of jobs
$26.84 - $28.63
4% of jobs
$28.63 - $30.42
3% of jobs
$30.42 - $32.21
1% of jobs
$12
$21
$32
How much do gpu repair jobs pay per hour?
What are the key skills and qualifications needed to thrive as a GPU Repair Technician, and why are they important?
What are GPU repair services?
Are GPU repairs worth IT?
What job designs GPUS?
How much does IT usually cost to repair a GPU?
What are some common challenges faced in GPU repair, and how can they be addressed?
What is the difference between Gpu Repair vs Gpu Technician?
| Aspect | Gpu Repair | Gpu Technician |
|---|---|---|
| Certifications | Hardware repair certifications, e.g., CompTIA A+ | Same as Gpu Repair, often includes electronics or hardware certifications |
| Work Environment | Repair shops, electronics labs, or service centers | Electronics labs, repair shops, or manufacturing facilities |
| Job Focus | Diagnosing and fixing GPU hardware issues | Diagnosing, repairing, and maintaining GPUs and related components |
| Industry Usage | Common in electronics repair industry | Used in electronics manufacturing and repair sectors |
Gpu Repair and Gpu Technician roles overlap significantly, focusing on diagnosing and fixing GPU hardware issues. Gpu Technicians often have broader responsibilities, including maintenance and testing, but both require similar certifications and work environments. The main difference lies in scope: Gpu Repair is more specialized in hardware fixes, while Gpu Technicians may handle a wider range of electronic components.
What jobs in the US pay 300,000 a year?

Full-time
This job post has expired today. Applications are no longer accepted.
Job description
Fluidstack is focused on building civilization-scale infrastructure for AI, aiming to deliver compute faster than anyone else. The Software Engineer, Compute (GPU) will own the health of the compute fleet, build automation for deployment and repair, and ensure the reliability and scalability of the GPU fleet.
Responsibilities:
• Own compute fleet health end to end. Build the metrics pipelines, alerting, and unified health view that tell you the true state of every GPU in production — across Kubernetes-orchestrated workloads and bare metal, at scale.
• Turn deployment/repair into a pipeline, not a procedure. Build and own the automation that takes a compute failure from detection through triage, parts management, and return to service. No one-off scripts, no heroics.
• Design and expand the GPU qualification platform. Burn-in, performance baselining, and NPI execution for every new GPU generation. You define what "good" looks like before hardware goes into production.
• Own Redfish and BMC tooling. Firmware-level telemetry, log collection at fleet scale, and the low-level access layer that repair automation and health tooling depend on.
• Own end-to-end reliability, scalability, and operation of the compute fleet at-scale. Fluidstack is building one of the largest GPU fleets in the world and that can only be accomplished with aggressive automation, tooling, and incident discipline.
Qualifications:
Required:
• You treat toil as a bug. Manual steps in a repair workflow are a backlog item, not a job description.
• You have an instinct for hardware. You're comfortable reasoning about failure modes at the firmware and silicon level, not just the software stack above it.
• You move toward ambiguity, not away from it. You walk into the fog, build the map, and explain it to everyone else.
• You learn at a steep slope. You reach real competence in an unfamiliar domain fast. We value this over existing expertise.
• You carry a pager without flinching. You run the incident, write the postmortem, fix the systemic cause, and move on.
• You're fluent with AI tooling. LLM APIs, MCP servers, and agentic frameworks, and you drive Claude Code, Cursor, or similar every day.
• You've shipped production automation that other teams depend on, and you're comfortable in any language using AI coding tools.
Preferred:
• Hardware lifecycle management and RMA automation.
• BMC/Redfish or IPMI tooling.
• GPU qualification or burn-in frameworks.
• Workflow and orchestration engines (Temporal, Cadence).
• Metrics and alerting pipelines (Prometheus, Grafana).
• Go or Python.
Company:
Fluidstack accelerates the world’s most ambitious AI projects by removing the bottlenecks to compute. Founded in 2017, the company is headquartered in London, GBR, with a team of 51-200 employees. The company is currently Growth Stage.