Neuveon

1 job near Columbus, OH

Site Reliability Engineer (SRE) GPU Infrastructure

Neuveon Inc

Santa Clara, CA • On-site

$66.50 - $88.25/hr

Other

Posted 4 days ago


Job description

We are seeking an experienced Site Reliability Engineer (SRE) for an hourly contract engagement supporting cutting-edge AI/GPU infrastructure. The consultant will troubleshoot complex Linux, server hardware, GPU, and infrastructure issues while collaborating with Engineering, Data Center Operations, and Capacity teams.

Key Responsibilities
  • Troubleshoot complex Linux, GPU, server hardware, firmware, and infrastructure issues.
  • Analyze system/kernel logs and BMC/Redfish telemetry to identify root causes.
  • Support hardware provisioning, repair, deployment, and production validation.
  • Develop automation, diagnostics, provisioning, and hardware repair tools.
  • Test and validate next-generation AI servers and GPU platforms.
  • Develop operational documentation, procedures, and best practices.
  • Collaborate with Hardware Engineering, Data Center Operations, and Capacity Planning teams.
  • Provide on-call and remote operational support as required.
Required Qualifications
  • Strong Linux administration and Linux internals knowledge.
  • Proven experience troubleshooting server hardware and infrastructure.
  • Strong understanding of GPU, hardware, firmware, networking, and systems troubleshooting.
  • Excellent root-cause analysis and problem-solving skills.
  • Experience with infrastructure provisioning and operations.
  • Strong communication and cross-functional collaboration skills.
  • Bachelor's degree in Computer Science or related field, or equivalent experience.
Nice to Have
  • Experience with large-scale GPU/AI infrastructure.
  • Programming experience with Python, Go, or similar languages.
  • Experience developing infrastructure or hardware automation tools.
  • Experience with BMC/Redfish and data center hardware operations.