1

Infiniband Jobs (NOW HIRING)

Practical knowledge of sophisticated networking for AI data centers, including InfiniBand and Ethernet fabric topologies, RDMA protocols, host networking and switches. * Experience with end-to-end ...

Practical knowledge of sophisticated networking for AI data centers, including InfiniBand and Ethernet fabric topologies, RDMA protocols, host networking and switches. * Experience with end-to-end ...

CPU and memory subsystems, PCIe devices, NVMe storage, Ethernet and InfiniBand networks. Investigate BIOS POST failures and boot sequence issues. Analyze BMC logs, sensor data, and firmware behavior.

AI Data Center Architect

Plano, TX · On-site

$61 - $78.50/hr

The ideal candidate will have strong experience with NVIDIA AI infrastructure, DGX/HGX, H200/GB200/Blackwell, InfiniBand, Spectrum-X, NVLink, Kubernetes, GPU scheduling, distributed training, and ...

New

Work with systems connected by various networking protocols, including Ethernet, NV Link and InfiniBand * Operate in Windows and Linux operating systems for test and debug tasks. * Collaborate with ...

Showing results 21-40

Infiniband information

See salary details

$11K

$83.3K

$113K

How much do infiniband jobs pay per year?

As of Aug 8, 2026, the average yearly pay for infiniband in the United States is $83,291.00, according to ZipRecruiter salary data. Most workers in this role earn between $35,500.00 and $108,000.00 per year, depending on experience, location, and employer.

What is InfiniBand?

Infiniband is a high-speed, low-latency networking technology commonly used in data centers and high-performance computing environments. It is designed to connect servers, storage systems, and network devices, providing much faster data transfer rates than traditional Ethernet. Infiniband supports scalable bandwidth and efficient communication, which makes it ideal for applications requiring rapid data movement, such as scientific simulations and large-scale database transactions. Its architecture also supports remote direct memory access (RDMA), which further reduces latency and CPU overhead.

What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?

To thrive as an InfiniBand Network Engineer, you need a strong background in computer networking, Linux system administration, and high-performance computing (HPC) environments, often supported by a degree in computer science or related field. Familiarity with InfiniBand architecture, experience with tools like OpenFabrics Enterprise Distribution (OFED), and certifications such as CompTIA Network+ are valuable. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for this role. These abilities are essential for ensuring efficient, reliable InfiniBand network performance in complex HPC or data center environments.

What are the typical responsibilities of an InfiniBand network engineer in a data center environment?

InfiniBand network engineers are primarily responsible for designing, deploying, and maintaining high-performance InfiniBand fabrics that connect servers and storage systems in data centers, especially in HPC (High-Performance Computing) environments. Their daily tasks include monitoring network performance, troubleshooting connectivity or latency issues, and performing firmware and driver updates on InfiniBand switches and host adapters. They also collaborate closely with system administrators and application teams to optimize throughput and ensure reliable, low-latency communication. Additionally, InfiniBand engineers often participate in capacity planning and help scale the network infrastructure to meet growing computational demands.

What is the difference between Infiniband vs Ethernet Network Engineer?

AspectInfinibandEthernet Network Engineer
Required CredentialsNetworking certifications, Cisco, Cisco CCNA, CCNPNetworking certifications, Cisco, CCNA, CCNP
Work EnvironmentData centers, high-performance computing environmentsCorporate networks, data centers, enterprise environments
Industry UsageHigh-performance computing, research institutionsBusiness, telecommunications, enterprise IT
Common Search/ComparisonYesYes

Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

More about Infiniband jobs
What cities are hiring for Infiniband jobs? Cities with the most Infiniband job openings:
What states have the most Infiniband jobs? States with the most job openings for Infiniband jobs include:
Infographic showing various Infiniband job openings in the United States as of August 2026, with employment types broken down into 97% Full Time, and 3% Contract. Highlights an 87% Physical, 2% Hybrid, and 11% Remote job distribution, with an average salary of $83,291 per year, or $40 per hour.

Sr. HPC Systems Engineer (IT@JH Research Computing)

ISACA

Baltimore, MD • On-site

$85.50 - $149.80/hr

Other

Posted 3 days ago

New


Job description

IT@JH Research Computing is seeking a Sr. HPC Systems Engineer who will design, build, and maintain advanced high-performance computing environments supporting Johns Hopkins University’s research mission. This position focuses on the reliable operation, configuration, and optimization of HPC and AI systems, including multi-node CPU and GPU clusters, high-speed InfiniBand and Ethernet networks, and large-scale parallel and object storage. The engineer implements and automates secure, efficient, and reproducible computing platforms used by faculty, researchers, and students across diverse scientific disciplines. Assignments include both ticket-based support and project-based deployments. The role operates with moderate independence, collaborating closely with the IT Architect, Research Computing, and reporting to the IT Manager for Research Computing to ensure scalable, sustainable, and high-performance systems that enable cutting‑edge scientific discovery.

Specific Duties & Responsibilities
  • Support and administer production systems used by researchers and Research Centers.
  • Provide technical leadership/project management for system configuration, implementation, management, and user support for both new and existing systems.
  • Research and recommend new functionality for HPC management and administration tools by exploring system-wide impacts, working with functional users to define current and future processes.
  • Expertise with architecting, operating, and debugging large scale HPC network and storage infrastructure, including MPI, NCCL, RDMA, Infiniband, and parallel file systems
  • Works with scientific support specialists and assigns tasks and provides oversight as appropriate to HPC engineering team to support scientific researchers who use a broad spectrum of applications from diverse fields.
  • Analyze results of server monitoring and implement changes to improve performance, processing, and utilization.
  • Propose, maintain, and enforce policies, practices and security procedures.
  • Provide break/fix support, setup/installation support, escalation support, and solutions support.
  • Collaborate closely with a variety of stakeholders, both internal and external, on all aspects of projects.
  • Other duties as assigned.
In Addition to the Duties Described Above
  • Deploy, configure, and maintain large-scale Linux-based HPC clusters comprising CPU and GPU nodes, high-speed interconnects, and parallel file systems.
  • Implement and optimize workload schedulers (Slurm) and job submission policies to maximize system throughput and fair-share usage.
  • Administer and monitor distributed storage systems (GPFS, Lustre, WekaFS, Ceph, MinIO) to ensure reliability and performance across multi-petabyte environments.
  • Maintain high-speed fabric and network infrastructure (Infiniband, Ethernet) to support low-latency data transfer and MPI workloads.
  • Support research groups in deploying, testing, and optimizing scientific applications and AI/ML workflows on shared computing resources.
  • Develop and maintain automation and monitoring frameworks for system provisioning, metrics collection, and alerting (Prometheus, Grafana, ELK).
  • Participate in capacity planning, hardware lifecycle management, and evaluation of new technologies in collaboration with architects and management.
  • Ensure security and compliance through configuration hardening, patch management, and integration with campus identity and access control systems.
  • Document system designs, procedures, and troubleshooting guides to support knowledge transfer and team continuity.
  • Contribute to a collaborative engineering culture that emphasizes service quality, innovation, and continuous improvement in research computing operations.
Minimum Qualifications
  • Bachelor’s degree.
  • Six years of related experience.
  • Additional education may substitute for required experience and additional related experience may substitute for required education beyond a high school diploma/graduation equivalent, to the extent permitted by the JHU equivalency formula.
Preferred Qualifications
  • Eight plus years of experience in high-performance computing systems administration or engineering, including experience with cluster management, workload scheduling (e.g., Slurm), and distributed or parallel storage.
  • Deep proficiency in Linux systems administration, configuration management (Ansible, Puppet, or Salt), performance monitoring, and tuning for HPC workloads.
  • Experience with high-speed interconnects (Infiniband, 100/400 Gb Ethernet) and parallel file systems (e.g., GPFS, Lustre, BeeGFS, or WekaFS).
  • Working knowledge of containerization and orchestration (Singularity, Docker, Kubernetes for HPC).
  • Ability to automate deployments and routine operations through scripting (Bash, Python).
  • Familiarity with data-center operations, GPU acceleration, and research software environments (e.g., CUDA, MPI, AI/ML frameworks).
  • Strong analytical and troubleshooting skills, with proven ability to support complex research workloads in multi-user, multi-tenant environments.
  • Experience collaborating with faculty and research groups to translate scientific requirements into practical and performant computing solutions.

Classified Title: Sr. HPC Systems Engineer
Role/Level/Range: ATP/04/PF
Starting Salary Range: $85,500 - $149,800 Annually (Commensurate w/exp.)
Employee group: Full Time
Schedule: Mon-Fri, 8:30am-5pm
FLSA Status:Exempt
Location: Johns Hopkins Bayview
Department name: IT@JH Research Computing
Personnel area: University Administration

#J-18808-Ljbffr