1

Nvidia Ai Infrastructure Jobs (NOW HIRING)

OR · On-site

$108K - $147K/yr

Joining NVIDIA's DGX Cloud AI Efficiency Team means contributing to the infrastructure that powers our innovative AI research. This team focuses on developing tools for optimizing efficiency and ...

Your day at NTT DATA The Senior Principal AI Infrastructure Architect is a highly skilled and ... Architect reference designs built on NVIDIA DGX/HGX SuperPOD, Dell AI Factory with NVIDIA, Cisco ...

The AI Hub team accelerates AI research by ensuring NVIDIA's AI infrastructure is used efficiently, transparently, and at scale. Our primary goal is to build a unified, self-service "single pane of ...

The AI Hub team accelerates AI research by ensuring NVIDIA's AI infrastructure is used efficiently, transparently, and at scale. Our primary goal is to build a unified, self-service "single pane of ...

next page

Showing results 1-20

Nvidia Ai Infrastructure information

See salary details

$80.5K

$154K

$198K

How much do nvidia ai infrastructure jobs pay per year?

As of Aug 10, 2026, the average yearly pay for nvidia ai infrastructure in the United States is $154,028.00, according to ZipRecruiter salary data. Most workers in this role earn between $113,000.00 and $197,000.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as an Nvidia AI Infrastructure Engineer, and why are they important?

To thrive as an Nvidia AI Infrastructure Engineer, you need a strong foundation in computer science, cloud computing, and distributed systems, often supported by a relevant degree and experience with large-scale AI workloads. Familiarity with Nvidia GPU technologies, CUDA programming, Kubernetes, and cloud platforms like AWS or Azure is typically required, along with certifications in cloud or AI infrastructure. Strong problem-solving skills, collaboration, and adaptability are essential soft skills for working across interdisciplinary teams and quickly evolving projects. These abilities ensure efficient deployment, scalability, and optimization of AI infrastructure, which are crucial for supporting advanced AI applications.

What is Nvidia AI Infrastructure?

Nvidia AI Infrastructure refers to the hardware, software, and cloud solutions provided by Nvidia to support the development, deployment, and scaling of artificial intelligence applications. This includes high-performance GPUs, networking technologies, data center platforms, and specialized software frameworks such as NVIDIA CUDA and NVIDIA AI Enterprise. Nvidia's AI infrastructure enables organizations to accelerate machine learning, deep learning, and data analytics workloads, both on-premises and in the cloud, delivering efficient and scalable AI solutions.

What are some typical challenges faced when managing AI infrastructure at Nvidia, and how can new team members prepare for them?

Managing AI infrastructure at Nvidia often involves supporting high-performance computing environments, scaling resources for large-scale machine learning workloads, and ensuring system reliability. New team members may face challenges such as optimizing GPU clusters, troubleshooting complex distributed systems, and staying current with rapidly evolving AI frameworks. To prepare, it's helpful to become familiar with Nvidia's hardware ecosystem, cloud-native technologies, and best practices for infrastructure automation. Proactively collaborating with software engineers, data scientists, and IT specialists is also essential for success in this dynamic environment.
Infographic showing various Nvidia Ai Infrastructure job openings in the United States as of August 2026, with employment types broken down into 74% Full Time, 22% Part Time, and 4% Contract. Highlights an 66% Physical, 3% Hybrid, and 31% Remote job distribution, with an average salary of $154,028 per year, or $74.1 per hour.

AI Operations & Infrastructure Engineer

Invictus International Consulting

Fort George G Meade, MD

$119K - $157K/yr

Full-time

Re-posted 18 days ago


Job description

Title: AI Operations & Infrastructure Engineer

Location: Fort Meade, MD

Clearance: TS/SCI with a CI Polygraph

Job Details:

  • Manage and maintain AI computing platforms, including GPUs and other specialized hardware
  • Install and configure GPU drivers and software
  • Oversee the AI software stack and tools
  • Implement and manage containerization technologies like Docker and Kubernetes
  • Configure and optimize networking infrastructure for AI workloads, including InfiniBand and Ethernet
  • Manage storage solutions for AI data, considering performance and capacity requirements
  • Deploy and manage data processing units (DPUs) to accelerate data center workloads
  • Monitor and manage AI cluster health and resource utilization
  • Implement workload management and scheduling tools like Slurm and Kubernetes
  • Ensure efficient power and cooling for AI infrastructure to maintain optimal operating conditions
  • Configure high-performance networking solutions for AI and machine learning workloads
  • Optimize network performance to ensure maximum throughput and minimal latency for AI computations
  • Implement and fine-tune network protocols to enhance data transfer speeds and efficiency
  • Integrate NVIDIA networking products with existing AI infrastructure, including servers, GPUs, and storage systems
  • Deploy networking solutions in data centers to ensure seamless connectivity between AI components
  • Diagnose and resolve networking issues impacting AI workloads to maintain optimal system performance
  • Provide technical support and guidance to teams managing AI infrastructure
  • Collaborate with data scientists, researchers, and IT professionals to understand networking requirements and challenges
  • Lead deployment and validation of servers and systems for AI enabled platforms
  • Configure and manage network topologies, BMC, OOB, TPM, power, and cooling
  • Install, upgrade, and validate GPU-based servers, BlueField DPUs, cables, and transceivers
  • Perform firmware upgrades, hardware validation, and storage setup
  • Configure and administer physical and logical resources, including M IG partitioning and BlueField platforms
  • Install and configure operating systems, cluster software, drivers, containers (Docker), and NGC CLI
  • Manage and orchestrate clusters using NVIDIA Base Command Manager, Slurm, Pyxis, Enroot, and Run: Ai
  • Perform stress, benchmarking, and burn-in tests using HPL, NCCL, NVIDIA Nemo, and ClusterKit
  • Verify cabling, firmware/software versions, and network signal quality
  • Troubleshoot and resolve hardware, software, storage, and performance faults
  • Replace faulty components and optimize systems for AMD/Intel platforms
  • Monitor, document, and report on cluster health, resource usage, and job performance
  • Ensure secure, efficient, and scalable operation of NVIDIA AI infrastructure, including user access and workload management

Requirements:

  • Qualified candidates must hold an active NVIDIA Professional Certification in either AI Networking, AI Infrastructure, or AI Operations
  • Prior direct, hands-on professional experience administering NVIDIA GPU and data processing unit (DPU) technologies, AI software stacks, and data center environments for high-performance AI workloads
  • Comprehensive expertise in deploying and maintaining AI compute platforms, requiring proficiency in containerization and workload orchestration using Docker, Kubernetes, Slurm, NVIDIA Base Command Manager, and Run:Ai
  • Must be capable of configuring physical and logical resources, including Multi-Instance GPU (MIG) partitioning and BlueField platforms, while overseeing critical facility elements such as power, cooling, and storage solutions
  • The ability to demonstrate advanced skills in AI networking, specifically configuring and optimizing high-performance InfiniBand and Ethernet fabrics to ensure maximum throughput and minimal latency
  • Current active TS/SCI clearance with a CI Polygraph

Equal Opportunity Employer/Veterans/Disabled