1

Infiniband Jobs in Texas (NOW HIRING)

Senior Deep Learning Communication Architect

Austin, TX · On-site

$128K - $174K/yr

Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...

Senior Solution Engineer, Networking

Austin, TX · On-site

$54.75 - $70.50/hr

They provide top support for high-speed interconnect technologies like InfiniBand, NVLink, and Spectrum-X that link GPUs and AI compute infrastructure. Candidates must have a software development ...

Key Responsibilities • Lead rack-and-stack of GPU servers, switches, and storage at colocation and modular data center sites • Cable compute, storage, and high-speed networking (InfiniBand and ...

AI Data Center Architect

Plano, TX · On-site

$61.25 - $78.75/hr

Design high-performance AI networking using InfiniBand, RoCE, Spectrum-X, NVLink, and BlueField DPUs. * Perform GPU cluster sizing, capacity planning, performance optimization, and scalability ...

Cloud Support Engineer

Dallas, TX · On-site

$55.25 - $73.75/hr

Understanding of HPC technologies such as Infiniband, RDMA, RoCE, and Software Defined Networking (SDN). Bonus Points: * Certifications: CKA, CKAD, CKS, KCNA, AWS Machine Learning - Specialty, Data ...

Experience in HPC technologies such as parallel/distributed files systems (e.g., Lustre, GPFS), high speed interconnect fabrics (e.g., Infiniband, Omni-Path), and HPC batch scheduling software suites ...

Experience in HPC technologies such as parallel/distributed files systems (e.g., Lustre, GPFS), high speed interconnect fabrics (e.g., Infiniband, Omni-Path), and HPC batch scheduling software suites ...

Understanding of HPC technologies such as Infiniband, RDMA, RoCE, and Software Defined Networking (SDN). Bonus Points: * Certifications: CKA, CKAD, CKS, KCNA, AWS Machine Learning - Specialty, Data ...

next page

Showing results 1-20

Infiniband information

What is InfiniBand?

Infiniband is a high-speed, low-latency networking technology commonly used in data centers and high-performance computing environments. It is designed to connect servers, storage systems, and network devices, providing much faster data transfer rates than traditional Ethernet. Infiniband supports scalable bandwidth and efficient communication, which makes it ideal for applications requiring rapid data movement, such as scientific simulations and large-scale database transactions. Its architecture also supports remote direct memory access (RDMA), which further reduces latency and CPU overhead.

What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?

To thrive as an InfiniBand Network Engineer, you need a strong background in computer networking, Linux system administration, and high-performance computing (HPC) environments, often supported by a degree in computer science or related field. Familiarity with InfiniBand architecture, experience with tools like OpenFabrics Enterprise Distribution (OFED), and certifications such as CompTIA Network+ are valuable. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for this role. These abilities are essential for ensuring efficient, reliable InfiniBand network performance in complex HPC or data center environments.

What are the typical responsibilities of an InfiniBand network engineer in a data center environment?

InfiniBand network engineers are primarily responsible for designing, deploying, and maintaining high-performance InfiniBand fabrics that connect servers and storage systems in data centers, especially in HPC (High-Performance Computing) environments. Their daily tasks include monitoring network performance, troubleshooting connectivity or latency issues, and performing firmware and driver updates on InfiniBand switches and host adapters. They also collaborate closely with system administrators and application teams to optimize throughput and ensure reliable, low-latency communication. Additionally, InfiniBand engineers often participate in capacity planning and help scale the network infrastructure to meet growing computational demands.

What is the difference between Infiniband vs Ethernet Network Engineer?

AspectInfinibandEthernet Network Engineer
Required CredentialsNetworking certifications, Cisco, Cisco CCNA, CCNPNetworking certifications, Cisco, CCNA, CCNP
Work EnvironmentData centers, high-performance computing environmentsCorporate networks, data centers, enterprise environments
Industry UsageHigh-performance computing, research institutionsBusiness, telecommunications, enterprise IT
Common Search/ComparisonYesYes

Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

What cities in Texas are hiring for Infiniband jobs? Cities in Texas with the most Infiniband job openings:
Infographic showing various Infiniband job openings in Texas as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% In-person job distribution.

Senior Deep Learning Communication Architect

NVIDIA

Austin, TX • On-site

$128K - $174K/yr

Full-time

Re-posted 22 hours ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 17 frontline employees who took The Breakroom Quiz

7th of 244 rated software companies


Job description

Job Summary:
NVIDIA is a leader in computer graphics and AI technology, seeking a Senior Deep Learning Communication Architect to enhance the performance of deep learning systems. The role involves optimizing communication protocols, collaborating with hardware and software teams, and exploring innovative communication technologies for distributed deep learning training and inference.
Responsibilities:
• The software architecture group at NVIDIA has openings for a Deep Learning Communication Architect. We scale the DNN models and training/inference frameworks to systems with hundreds of thousands of nodes.
• Optimizing communication performance: Identify and eliminate bottlenecks in data transfer and synchronization during distributed deep learning training and inference.
• Designing efficient communication protocols: Develop and implement communication algorithms and protocols tailored for deep learning workloads, minimizing communication overhead and latency.
• Hardware and software co-craft: Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC, NVSHMEM).
• Exploring innovative communication technologies: Research and evaluate new communication technologies and techniques to enhance the performance and scalability of deep learning systems.
• Developing and implementing solutions: Build proofs-of-concept, conduct experiments, and perform quantitative modeling to validate and deploy new communication strategies.
Qualifications:
Required:
• A Ph.D., Masters, or BS in Computer Science (CS), Electrical Engineering (EE), Computer Science and Electrical Engineering (CSEE), or a closely related field or equivalent experience.
• 6+ years of experience in Building DNNs, Scaling of DNNs, Parallelism of DNN frameworks, or deep learning training and inference workloads.
• Experience in evaluating, analyzing, and optimizing LLM training and inference performance of state-of-the-art models on cutting-edge hardware.
• Deep understanding of parallelism techniques, including Data Parallelism, Pipeline Parallelism, Tensor Parallelism, Expert Parallelism, and FSDP.
• Understanding of the emerging serving architectures like Disaggregated Serving and inference servers like Dynamo and Triton.
• Proficiency in developing code for one or more deep neural network (DNN) training and Inference frameworks, such as PyTorch, TensorRT-LLM, vLLM, SGLang.
• Strong programming skills in C++ and Python.
• Familiarity with GPU computing, including CUDA and OpenCL, and familiarity with InfiniBand and RoCE networks.
Preferred:
• Prior contributions to one or more DNN training and Inference frameworks as part of your previous work experience.
• Deep understanding and contributions to the scaling of LLMs on large-scale systems.
Company:
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US

Year founded

1993