1

Nvidia Ai Infrastructure Jobs (NOW HIRING)

Senior DGX Cloud AI Infrastructure Software Engineer

OR ยท On-site +1

$108K - $147K/yr

Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and ...

Senior DGX Cloud AI Infrastructure Software Engineer

OR ยท On-site +1

$108K - $147K/yr

Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and ...

next page

Showing results 1-20

Nvidia Ai Infrastructure information

See salary details

$80.5K

$154K

$198K

How much do nvidia ai infrastructure jobs pay per year?

As of Sep 5, 2026, the average yearly pay for nvidia ai infrastructure in the United States is $154,028.00, according to ZipRecruiter salary data. Most workers in this role earn between $113,000.00 and $197,000.00 per year, depending on experience, location, and employer.

What is Nvidia AI Infrastructure?

Nvidia AI Infrastructure refers to the hardware, software, and cloud solutions provided by Nvidia to support the development, deployment, and scaling of artificial intelligence applications. This includes high-performance GPUs, networking technologies, data center platforms, and specialized software frameworks such as NVIDIA CUDA and NVIDIA AI Enterprise. Nvidia's AI infrastructure enables organizations to accelerate machine learning, deep learning, and data analytics workloads, both on-premises and in the cloud, delivering efficient and scalable AI solutions.

What are some typical challenges faced when managing AI infrastructure at Nvidia, and how can new team members prepare for them?

Managing AI infrastructure at Nvidia often involves supporting high-performance computing environments, scaling resources for large-scale machine learning workloads, and ensuring system reliability. New team members may face challenges such as optimizing GPU clusters, troubleshooting complex distributed systems, and staying current with rapidly evolving AI frameworks. To prepare, it's helpful to become familiar with Nvidia's hardware ecosystem, cloud-native technologies, and best practices for infrastructure automation. Proactively collaborating with software engineers, data scientists, and IT specialists is also essential for success in this dynamic environment.

What are the key skills and qualifications needed to thrive as an Nvidia AI Infrastructure Engineer, and why are they important?

To thrive as an Nvidia AI Infrastructure Engineer, you need a strong foundation in computer science, cloud computing, and distributed systems, often supported by a relevant degree and experience with large-scale AI workloads. Familiarity with Nvidia GPU technologies, CUDA programming, Kubernetes, and cloud platforms like AWS or Azure is typically required, along with certifications in cloud or AI infrastructure. Strong problem-solving skills, collaboration, and adaptability are essential soft skills for working across interdisciplinary teams and quickly evolving projects. These abilities ensure efficient deployment, scalability, and optimization of AI infrastructure, which are crucial for supporting advanced AI applications.
Infographic showing various Nvidia Ai Infrastructure job openings in the United States as of August 2026, with employment types broken down into 76% Full Time, 20% Part Time, and 4% Contract. Highlights an 66% Physical, 4% Hybrid, and 30% Remote job distribution, with an average salary of $154,028 per year, or $74.1 per hour.

NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems)

Catapult Solutions Group

New York, NY โ€ข On-site

$125/hr

Contractor

Medical, Retirement, PTO

This job post hasย expired 2 days ago.ย Applications are no longer accepted.


Job description

NVIDIA AI Infrastructure & Kubernetes Platform Engineer (DGX Systems)
Department: Infrastructure Engineering
Location / Remote Policy: Remote
Role Type: Contract - 6-month initial engagement
About Our Client
Our client is a technology and professional services firm founded in 2015 on the strength of its founders' 30 years of industry experience. They set out to bridge a gap in professional services - to be a true partner rather than just a vendor - delivering expert guidance, innovative solutions, and personalized service at a cost-effective rate. Their mission is to empower businesses to succeed in the digital era, harnessing technology to drive transformation, innovation, and growth. Guided by a "make a customer, not a sale" philosophy, they lead with a customer-first approach and a team of senior-level engineers sourced from the world's leading OEMs, including AWS, Palo Alto Networks, Cisco, and Microsoft.
Job Description
Our client is seeking a highly skilled AI Infrastructure & Kubernetes Platform Engineer with a proven track record deploying and managing NVIDIA DGX-based AI clusters, orchestrating containerized AI workloads on Kubernetes, and ensuring secure, high-throughput operations across InfiniBand-powered networks. You'll bring a strong certification foundation across both Kubernetes (CKA, CKAD, CKS) and NVIDIA's AI infrastructure stack, paired with hands-on experience across DGX, BlueField, and high-speed networking.
This role is central to supporting AI/ML infrastructure at scale - enabling efficient training and inference for complex models and integrating NVIDIA's compute, storage, and fabric solutions with modern DevOps practices. Day to day, you'll own DGX cluster operations, architect GPU-accelerated Kubernetes platforms, tune InfiniBand fabric for throughput, and harden the environment through DPU-enhanced security.
You'll work at the intersection of infrastructure, DevOps, and AI/ML, keeping the platform reliable and cost-efficient for the teams that depend on it. The ideal candidate is deeply hands-on, obsessed with performance and security, and energized by operating some of the most advanced AI compute available.
Duties and Responsibilities
AI Infrastructure Operations
  • Deploy and manage NVIDIA DGX BasePODs and SuperPODs for high-performance AI workloads.
  • Oversee DGX system lifecycle operations, including provisioning, monitoring, firmware upgrades, and capacity planning.
  • Operate Base Command Manager to manage GPU clusters, schedule workloads, and integrate with MLOps tools.
  • Perform DGX node health validation, NCCL interconnect testing, and NVLink topology verification after deployments or hardware changes.

Kubernetes Platform Engineering
  • Architect secure, scalable Kubernetes clusters optimized for GPU-accelerated workloads using the NVIDIA GPU Operator.
  • Apply CKA/CKAD/CKS expertise to develop, deploy, and secure AI applications on Kubernetes.
  • Implement CI/CD pipelines and GitOps methodologies for deploying and managing ML workflows.

High-Performance Networking & DPUs
  • Administer InfiniBand networks and BlueField DPUs using Unified Fabric Manager (UFM).
  • Enable NVLink/NVSwitch performance across GPU nodes and tune fabric configurations for minimal latency and maximum throughput.
  • Use BlueField to offload storage, firewalling, and telemetry, strengthening AI workload security and performance.

Security & Compliance
  • Apply CKS best practices to secure containerized AI environments.
  • Configure runtime security, secrets management, network segmentation, and auditing across DPU-enhanced Kubernetes deployments.
  • Support zero-trust initiatives by enforcing workload identity, RBAC policies, and supply-chain integrity across AI container images and model artifacts.

Monitoring, Telemetry & Optimization
  • Monitor GPU, CPU, and I/O performance using NVIDIA DCGM, Prometheus, Grafana, and Base Command APIs.
  • Tune system performance and model-training pipelines for cost-efficiency and throughput.
  • Build and maintain operational runbooks, incident-response playbooks, and SLA dashboards covering GPU utilization, thermal thresholds, and fabric health.

Required Experience/Skills
Certifications
  • Certified Kubernetes Administrator (CKA)
  • Certified Kubernetes Application Developer (CKAD)
  • Certified Kubernetes Security Specialist (CKS)
  • NVIDIA Certified Associate: AI Infrastructure & Operations (NCA-AIIO)
  • NVIDIA Certified Professional: AI Infrastructure (NCP-AII)
  • NVIDIA Certified Professional: AI Operations (NCP-AIO)
  • NVIDIA Certified Professional: AI Networking (NCP-AIN)

Hands-On Expertise
  • DGX System, BasePOD, and SuperPOD administration
  • BlueField DPU configuration and operations
  • InfiniBand fabric and UFM management
  • Base Command Manager for workload orchestration

Technical Skills
  • Kubernetes, Helm, and the NVIDIA GPU Operator
  • DevOps tooling: Ansible, Terraform, GitOps, CI/CD pipelines
  • Programming/scripting: Python, YAML, Bash

Nice-to-Haves
  • Kubeflow and broader MLOps pipeline experience.
  • Parallel/HPC storage: NFS, BeeGFS, Lustre.
  • Advanced networking: RoCE, RDMA, gRPC, and DPU offload tuning.

Education
Bachelor's degree in Computer Science, Engineering, or a related field - or equivalent hands-on experience.
Pay & Benefits Summary
  • Pay Rate: $125/hr
  • Benefits: Eligible for a comprehensive benefits package, including medical/health coverage, paid time off, and 401(k) retirement savings.

Call-to-Action
Operate the cutting edge of AI compute. Apply today and put your DGX, Kubernetes, and NVIDIA expertise to work.
Keywords: NVIDIA DGX | SuperPOD | Kubernetes | CKA / CKAD / CKS | NCA-AIIO | NCP-AII | NCP-AIO | NCP-AIN | InfiniBand | BlueField DPU | UFM | GPU Operator | Kubeflow | RDMA | RoCE | MLOps