1

Gpu Computing System Software Manager Jobs (NOW HIRING)

Senior System Engineer - GPU Platforms

San Jose, CA ยท On-site

$122K - $167K/yr

Collaborate with Architecture, Systems, Software, Validation, Product Management, other engineering ... Experience with GPU computing, accelerators, or comparable high-performance computing technologies

Sr. System Engineer/GPU Platforms

San Jose, CA ยท On-site

$123K - $169K/yr

Collaborate with internal Architecture, Systems, Software, Validation, Product Management, and ... Experience with GPU computing, accelerators, or comparable high-performance computing technologies.

Automotive, VR, Gaming, Deep Learning, and High Performance Computing. See your efforts in action ... As a system software engineer in the Developer Tools group, you will be developing software that ...

next page

Showing results 1-20

Gpu Computing System Software Manager information

See salary details

$139.5K

$155K

$169K

How much do gpu computing system software manager jobs pay per year?

As of Sep 12, 2026, the average yearly pay for gpu computing system software manager in the United States is $155,000.00, according to ZipRecruiter salary data. Most workers in this role earn between $147,500.00 and $162,500.00 per year, depending on experience, location, and employer.

What are popular job titles related to Gpu Computing System Software Manager jobs?

For Gpu Computing System Software Manager jobs, the most frequently searched job titles are:

Infographic showing various Gpu Computing System Software Manager job openings in the United States as of August 2026, with employment types broken down into 88% Full Time, 11% Part Time, and 1% Contract. Highlights an 80% Physical, 2% Hybrid, and 18% Remote job distribution, with an average salary of $155,000 per year, or $74.5 per hour.

Senior System Engineer - GPU Platforms

San Jose, CA โ€ข On-site

$122K - $167K/yr

Other

Posted 6 days ago


Job description

  • Support system bring-up, configuration, integration, validation, and troubleshooting of advanced GPU server platforms
  • Execute and support GPU platform qualification activities, including NVIDIA NVQUAL or equivalent validation processes
  • Install, configure, and troubleshoot Linux, GPU drivers, CUDA environments, firmware, libraries, and related software components
  • Diagnose complex system issues using logs, telemetry, diagnostics, and vendor tools, and drive issues to resolution or appropriate engineering escalation
  • Support multi-GPU server platforms throughout qualification, product launch, and post-release engineering activities
  • Participate in customer-facing POC/EVAL engagements, including system preparation, technical calls, debugging, and issue resolution
  • Collaborate with Architecture, Systems, Software, Validation, Product Management, other engineering teams, and external technology partners
  • Develop technical documentation, troubleshooting guides, and best practices
  • Deliver technical presentations, training sessions, and internal knowledge-sharing activities
  • Serve as a technical resource and mentor for other engineers when appropriate
Requirements
  • Bachelorโ€™s degree in Computer Engineering, Electrical Engineering, Computer Science, Information Technology, or a related discipline, or equivalent practical experience
  • 5โ€“15 years of relevant industry experience in systems engineering, server engineering, platform engineering, validation, technical enablement, HPC, AI infrastructure, or a related field
  • Strong knowledge of enterprise server hardware and system architecture
  • Hands-on experience with Linux server environments
  • Experience installing, configuring, validating, and troubleshooting server hardware and software
  • Strong system-level troubleshooting and root-cause-analysis skills
  • Working knowledge of PCIe architectures and high-performance I/O
  • Experience with GPU computing, accelerators, or comparable high-performance computing technologies
  • Ability to independently manage complex technical assignments and drive issues toward resolution
  • Strong written and verbal communication skills
  • Ability to work effectively with cross-functional and geographically distributed engineering teams
  • Comfortable participating in customer-facing technical discussions
  • Preferred: Hands-on experience with NVIDIA data center or professional GPU platforms
  • Preferred: Experience with CUDA and NVIDIA GPU software environments
  • Preferred: Experience with NVIDIA NVQUAL or similar platform qualification processes
  • Preferred: Experience with 4-GPU or 8-GPU server platforms
  • Preferred: Familiarity with NVIDIA Blackwell, B200, Rubin, or comparable accelerator architectures
  • Preferred: Knowledge of PCIe topology, NUMA, DMA, IOMMU, and GPU-to-NIC communication
  • Preferred: Experience with GPUDirect RDMA, InfiniBand, RoCE, or high-speed Ethernet
  • Preferred: Familiarity with NCCL, NVML, DCGM, Fabric Manager, or similar GPU diagnostic and management tools
  • Preferred: Experience with Docker, containers, Kubernetes, or related orchestration technologies
  • Preferred: Experience supporting AI, machine learning, HPC, or accelerated computing environments
  • Preferred: Experience with customer POCs, technical evaluations, or engineering escalations
  • Preferred: Experience delivering technical training or knowledge-sharing sessions
  • Bash, Python, or other scripting experience is a plus
Core Competencies

Demonstrates expertise in GPU server platform support, including installation, configuration, and troubleshooting of Linux environments and GPU technologies. Proficient in system-level diagnostics, technical documentation, and cross-functional collaboration to drive complex technical assignments to resolution.

Highest-signal resume keywords
  • GPU Computing
  • Linux Server Environments
  • System-Level Troubleshooting
  • NVIDIA NVQUAL
  • Technical Documentation
Hard Skills
  • GPU Drivers
  • CUDA
  • Server Hardware Configuration
  • Root-Cause Analysis
  • PCIe Architectures
  • High-Performance Computing
  • NVIDIA Data Center Platforms
  • Docker
  • Python Scripting
  • Bash Scripting
Soft Skills
  • Strong Communication Skills
  • Cross-Functional Collaboration
  • Customer-Facing Technical Discussions
  • Mentoring
Industry Keywords
  • Systems Engineering
  • Server Engineering
  • Platform Engineering
  • Technical Enablement
  • AI Infrastructure
  • HPC
Tools & Technologies
  • NCCL
  • NVML
  • DCGM
  • Fabric Manager
  • InfiniBand
  • RoCE
  • High-Speed Ethernet
  • Kubernetes
#J-18808-Ljbffr