1

Gpu Repair Jobs (NOW HIRING)

Principal Software Engineer, GPU Compute

San Mateo, CA · On-site

$153K - $206K/yr

Drive GPU reliability and performance at fleet scale, defining the detection, diagnosis, and automated repair of unhealthy accelerators before they impact production. * Evaluate and onboard new GPU ...

As an Debug Repair Technician you will: * Perform advanced board-level troubleshooting and ... GPU, FPGA, ASIC, memory, and power subsystems Diagnose high speed digital, power distribution, and ...

Responsibilities : • Operate and support GPU workflows across RMA, Service, Repair, TS, Purchasing, System PM, and Production environments • Analyze GPU and system logs using CLI tools • ...

Associate Product Manager

San Jose, CA · On-site

$90K - $105K/yr

The Associate Product Manager will support and coordinate GPU-related operations across RMA, Service, Repair, Technical Support, and Production teams, with a focus on improving workflow efficiency ...

Build the repair pipeline that keeps pace with a fleet of 10s to 100s of GWs: at our scale, a GPU failure isn't a ticket. It's a throughput problem. We're building the automation that takes a chip ...

The role Nebius operates large-scale, GPU-dense AI infrastructure across mission-critical data ... You will lead on-site rack bring-up, validate NVIDIA-based AI systems, coordinate repairs, and ...

next page

Showing results 1-20

Gpu Repair information

See salary details

$12

$21

$32

How much do gpu repair jobs pay per hour?

As of Jul 22, 2026, the average hourly pay for gpu repair in the United States is $21.41, according to ZipRecruiter salary data. Most workers in this role earn between $17.55 and $24.04 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a GPU Repair Technician, and why are they important?

To thrive as a GPU Repair Technician, you need a strong knowledge of computer hardware, electronics troubleshooting, and experience with soldering and component-level repair, often supported by relevant technical certifications or training. Familiarity with diagnostic tools, multimeters, oscilloscopes, and GPU testing software is typically required. Attention to detail, problem-solving skills, and effective communication help technicians accurately diagnose issues and explain solutions to clients. These skills ensure reliable GPU repairs, minimize downtime, and maintain customer trust in high-performance computing environments.

What are GPU repair services?

GPU repair services involve diagnosing and fixing issues with graphics processing units (GPUs) used in computers and gaming systems. Common problems include overheating, faulty memory chips, broken fans, or damaged solder joints. Technicians may clean the GPU, replace faulty components, or reflow/reball solder connections to restore functionality. Proper diagnostics are essential to determine whether repair or full replacement is more cost-effective. Reliable GPU repair can extend the device's lifespan and save money compared to purchasing a new graphics card.

Are GPU repairs worth IT?

GPU repair technicians diagnose and fix issues with graphics cards, often involving troubleshooting hardware components and soldering skills. The job can be worthwhile for those with technical expertise, as demand exists for repairing or refurbishing GPUs, especially in gaming and professional markets. However, the complexity of modern GPUs and the cost of repairs can impact the profitability and feasibility of repair jobs.

What job designs GPUS?

A GPU repair technician or engineer is responsible for designing, diagnosing, and repairing graphics processing units. They often work with hardware schematics, diagnostic tools, and may need knowledge of electronics, soldering, and computer architecture. Certification in electronics or computer hardware can be beneficial for this role.

How much does IT usually cost to repair a GPU?

For GPU repair technicians, the cost to repair a GPU typically ranges from $100 to $300, depending on the extent of damage and parts needed. Common repairs include fixing overheating issues, replacing damaged components, or reballing the GPU, often requiring specialized tools and diagnostic skills.

What are some common challenges faced in GPU repair, and how can they be addressed?

Technicians working in GPU repair often encounter challenges such as diagnosing subtle hardware faults, dealing with delicate soldering work, and sourcing replacement components for discontinued models. Addressing these issues requires strong troubleshooting skills, familiarity with GPU architecture, and access to specialized tools like hot air rework stations and microscopes. Collaboration with other repair technicians and staying updated on manufacturer guidelines are also essential for successful repairs and ongoing professional development.

What is the difference between Gpu Repair vs Gpu Technician?

AspectGpu RepairGpu Technician
CertificationsHardware repair certifications, e.g., CompTIA A+Same as Gpu Repair, often includes electronics or hardware certifications
Work EnvironmentRepair shops, electronics labs, or service centersElectronics labs, repair shops, or manufacturing facilities
Job FocusDiagnosing and fixing GPU hardware issuesDiagnosing, repairing, and maintaining GPUs and related components
Industry UsageCommon in electronics repair industryUsed in electronics manufacturing and repair sectors

Gpu Repair and Gpu Technician roles overlap significantly, focusing on diagnosing and fixing GPU hardware issues. Gpu Technicians often have broader responsibilities, including maintenance and testing, but both require similar certifications and work environments. The main difference lies in scope: Gpu Repair is more specialized in hardware fixes, while Gpu Technicians may handle a wider range of electronic components.

What jobs in the US pay 300,000 a year?

In the US, high-paying roles related to GPU repair or computer hardware often include senior hardware engineers, IT directors, or specialized technical consultants, typically requiring advanced skills, certifications, and experience. These positions may reach or exceed $300,000 annually, especially in management or consulting roles within technology companies. However, most GPU repair technicians or similar technical roles tend to have lower salaries unless combined with managerial responsibilities or specialized expertise.
More about Gpu Repair jobs
What cities are hiring for Gpu Repair jobs? Cities with the most Gpu Repair job openings:
What states have the most Gpu Repair jobs? States with the most job openings for Gpu Repair jobs include:
Infographic showing various Gpu Repair job openings in the United States as of July 2026, with employment types broken down into 1% Locum Tenens, 61% Full Time, 1% Part Time, 1% Contract, and 36% Nights. Highlights an 85% Physical, 6% Hybrid, and 9% Remote job distribution, with an average salary of $44,528 per year, or $21.4 per hour.

Senior Staff Data Center Operations Engineer, GPU Hardware Architecture

Crusoe

Sunnyvale, CA • On-site

$129K - $173K/yr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Posted 16 days ago


Job description

Crusoe is on a mission to accelerate the abundance of energy and intelligence. As the only vertically integrated AI infrastructure company built from the ground up, we own and operate each layer of the stack — from electrons to tokens — to power the world's most ambitious AI workloads. When you join Crusoe, you join a team that is building the future, faster.

We're in the midst of the greatest industrial revolution of our time. The demand for AI compute is boundless, and power is a bottleneck. We're solving that — with an energy-first approach that makes AI infrastructure better for the world and faster for the people innovating with AI.

We're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved — people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services.

If you want to do the most meaningful work of your career, help our customers and partners advance their AI strategies, and be part of a high-performing team that believes in each other, come build with us at Crusoe.

The Mission

Crusoe is building the world’s most climate-aligned AI infrastructure. As we scale toward unprecedented power densities and liquid-cooled architectures, the gap between "Data Center Design" and "Silicon Reality" must be bridged.

We are seeking a Senior Staff Data Center Operations Engineer, GPU Hardware Architecture to be the definitive technical authority on GPU platforms within the Data Center Engineering and Operations organization. Your mission is twofold: act as the primary technical consultant to our Data Center Engineering team to ensure future facilities are built for next-gen silicon, and provide the Operations team with the specialized tooling, SOPs, and predictive strategies needed to maintain peak cluster health.

 
The Strategic Bridge
  • For DC Engineering: You are the internal consultant. You translate upcoming GPU power/thermal roadmaps (NVIDIA/AMD) into design requirements for our next-generation facilities.

  • For Site Operations: You are the "Technical Enabler." You develop the diagnostic tools and technical SOPs that enable field technicians to resolve complex GPU issues with surgical accuracy.

  • For Sourcing: You are the "Technical Strategist." You define the technical sparing requirements and site-level inventory needs based on hardware failure telemetry.

 
Key Responsibilities
  • Engineering Education & Design Support: Provide deep-dive technical guidance to the Data Center Engineering team on upcoming silicon (e.g., NVIDIA Blackwell/Rubin, AMD MI350/400). Ensure future facility designs for power, cooling, and rack-spacing are ready for 2000W+ per-chip densities.

  • Predictive Operations & Telemetry: Leverage AI/ML methodologies to analyze fleet-wide telemetry (power draws, thermal gradients, and error rates). You will lead the transition from reactive troubleshooting to predictive maintenance, identifying "pre-failure" patterns in HBM or NVLink components before they impact customer training runs.

  • Technical Sparing Architecture: Architect the site-level sparing strategy from a technical perspective. Use failure telemetry and MTBF data to define the "Critical Spares List" and stocking levels required at each site to meet cluster uptime targets, providing these requirements to Sourcing for execution.

  • Operational Tooling & SOPs: Build the "Operational Blueprint" for the field. Create precision SOPs for high-stakes GPU repairs (e.g., baseboard swaps, manifold maintenance) and develop diagnostic tooling that allows Site Ops to identify NVLink flapping, PCIe degradations, or thermal throttling.

  • Advanced Troubleshooting & RCA: Act as the Tier-3 escalation point for the most complex hardware failures in the production environment. Lead Root Cause Analysis (RCA) on systemic issues that span the boundary between hardware and facility environmental factors.

  • Silicon Roadmap Authority: Maintain a 24-month forward-looking view of NVIDIA and AMD architectures. Educate internal stakeholders on how transitions in HBM4, interconnect speeds, and liquid-cooling will impact Crusoe’s physical infrastructure.

  • Vendor & VAR Technical Lead: Support the technical relationship with OEMs and VARs. Audit their hardware builds, review their technical bulletins, and ensure their hardware roadmaps align with Crusoe’s operational and engineering standards.

 
Technical Requirements
  • Silicon & Fabric Mastery: Expert-level knowledge of NVIDIA (Hopper/Blackwell/Rubin) and AMD (Instinct) architectures. Mastery of the physical and logical layers of NVLink, NVSwitch, and InfiniBand.

  • Infrastructure Bridge-Building: Ability to translate "Silicon Data Sheets" into "Mechanical Engineering Requirements." You can explain how a GPU's specific heat-load profile affects CDU sizing and secondary loop design.

  • Data-Driven Diagnostics: Proficient in Python, Go, or Bash to build telemetry and health-check tools (utilizing DCGM and ROCm). Experience using large datasets or basic ML frameworks to build "Smart Monitoring" that filters critical health signals from noise.

  • Operational Reliability Analysis: Experience using failure telemetry to inform site-level sparing requirements and field-service workflows.

  • Thermal Management: Deep understanding of the operational realities of Direct-to-Chip (D2C) cooling, including fluid dynamics, pressure-drop curves, and the lifecycle of dripless couplings.

 
Qualifications
  • 10+ years in Hardware Engineering, Systems Architecture, or Data Center Infrastructure.

  • The "Consultant" Mindset: Proven track record of educating and influencing cross-functional teams (specifically Engineering and Operations).

  • GPU Authority: You have managed or architected GPU clusters at scale (thousands of nodes) at a hyperscaler, a GPU-specialized cloud, or a major silicon vendor.

Education: B.S. or M.S. in Electrical Engineering, Computer Engineering, or a related technical field.

 

Benefits:

  • Competitive compensation

  • Restricted Stock Units

  • Paid time off & paid holidays

  • Comprehensive health, dental & vision insurance

  • Employer contributions to HSA account

  • Paid parental leave

  • Paid life insurance, short-term and long-term disability

  • Professional development & tuition reimbursement

  • Mental health & wellness support

  • Commuter benefits (parking & transit)

  • Cell phone stipend

  • 401(k) Retirement plan with company match up to 4% of salary

  • Volunteer time off

 

Compensation Range

Compensation will be paid in the range of up to $179,000 -$218,000 + Bonus. Restricted Stock Units are included in all offers. Compensation to be determined by the applicants knowledge, education, and abilities, as well as internal equity and alignment with market data.

Crusoe is an Equal Opportunity Employer. Employment decisions are made without regard to race, color, religion, disability, genetic information, pregnancy, citizenship, marital status, sex/gender, sexual preference/ orientation, gender identity, age, veteran status, national origin, or any other status protected by law or regulation.