1

Infiniband Jobs in Texas (NOW HIRING)

Forward Deployed Engineer

Dallas, TX · On-site

$150 - $190/hr

Mastery of high-speed networking, including InfiniBand, RoCE, and 100GbE+ * Ability to debug complex routing, latency, and throughput issues in distributed environments * Proficiency in Python, Go ...

New

Principal Network Engineer

Houston, TX · On-site

$180 - $270/hr

Designing, reviewing, and evolving large-scale Infiniband and RoCE fabric architectures to support future growth and workload demands * Acting as the senior escalation point for the most complex ...

New

Principal Network Engineer

Houston, TX · On-site

$180 - $240/hr

Designing, reviewing, and evolving large-scale Infiniband and RoCE fabric architectures to support future growth and workload demands * Acting as the senior escalation point for the most complex ...

New

Senior Solution Engineer, Networking

Austin, TX

$54.75 - $70.50/hr

They provide top support for high-speed interconnect technologies like InfiniBand, NVLink, and Spectrum-X that link GPUs and AI compute infrastructure. Candidates must have a software development ...

Senior Deep Learning Communication Architect

Austin, TX · On-site

$128K - $174K/yr

Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...

Senior GPU Systems & Fabric Engineer

Austin, TX · On-site +1

$103K - $141K/yr

Configure and optimize high-performance host networking stacks, including RDMA, SR-IOV, RoCEv2, and InfiniBand, ensuring line-rate throughput for distributed AI training. * Build and manage automated ...

Posted today

Senior AI Storage Infrastructure Engineer

Austin, TX · On-site +1

$107K - $146K/yr

Collaborate with the GPU Systems & Fabric team to ensure the storage layer is fully optimized for RDMA and high-speed interconnects (InfiniBand, RoCE). * Implement automated monitoring and alerting ...

Posted today

Key Responsibilities • Lead rack-and-stack of GPU servers, switches, and storage at colocation and modular data center sites • Cable compute, storage, and high-speed networking (InfiniBand and ...

next page

Showing results 1-20

Infiniband information

What is InfiniBand?

Infiniband is a high-speed, low-latency networking technology commonly used in data centers and high-performance computing environments. It is designed to connect servers, storage systems, and network devices, providing much faster data transfer rates than traditional Ethernet. Infiniband supports scalable bandwidth and efficient communication, which makes it ideal for applications requiring rapid data movement, such as scientific simulations and large-scale database transactions. Its architecture also supports remote direct memory access (RDMA), which further reduces latency and CPU overhead.

What are the typical responsibilities of an InfiniBand network engineer in a data center environment?

InfiniBand network engineers are primarily responsible for designing, deploying, and maintaining high-performance InfiniBand fabrics that connect servers and storage systems in data centers, especially in HPC (High-Performance Computing) environments. Their daily tasks include monitoring network performance, troubleshooting connectivity or latency issues, and performing firmware and driver updates on InfiniBand switches and host adapters. They also collaborate closely with system administrators and application teams to optimize throughput and ensure reliable, low-latency communication. Additionally, InfiniBand engineers often participate in capacity planning and help scale the network infrastructure to meet growing computational demands.

What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?

To thrive as an InfiniBand Network Engineer, you need a strong background in computer networking, Linux system administration, and high-performance computing (HPC) environments, often supported by a degree in computer science or related field. Familiarity with InfiniBand architecture, experience with tools like OpenFabrics Enterprise Distribution (OFED), and certifications such as CompTIA Network+ are valuable. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for this role. These abilities are essential for ensuring efficient, reliable InfiniBand network performance in complex HPC or data center environments.

What is the difference between Infiniband vs Ethernet Network Engineer?

AspectInfinibandEthernet Network Engineer
Required CredentialsNetworking certifications, Cisco, Cisco CCNA, CCNPNetworking certifications, Cisco, CCNA, CCNP
Work EnvironmentData centers, high-performance computing environmentsCorporate networks, data centers, enterprise environments
Industry UsageHigh-performance computing, research institutionsBusiness, telecommunications, enterprise IT
Common Search/ComparisonYesYes

Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

What job categories do people searching Infiniband jobs in Texas look for?

The top searched job categories for Infiniband jobs in Texas are:

What cities in Texas are hiring for Infiniband jobs?

Cities in Texas with the most Infiniband job openings:

Infographic showing various Infiniband job openings in Texas as of August 2026, with employment types broken down into 98% Full Time, and 2% Contract. Highlights an 81% Physical, 4% Hybrid, and 15% Remote job distribution.

GPU DC East-West Network SRE Expert (SME)

BitDeer

Austin, TX • On-site, Remote

Full-time

Posted 9 days ago


Job description

Bitdeer is a world-leading technology company for AI and Bitcoin mining infrastructure.
Bitdeer is committed to providing comprehensive Bitcoin mining solutions for its customers and building AI computational infrastructure to support the AI revolution. Bitdeer handles complex processes involved in computing such as equipment procurement, transport logistics, data center design and construction, equipment management, and daily operations. Bitdeer also offers advanced cloud capabilities to customers with high demand for artificial intelligence.
Headquartered in Singapore, Bitdeer has deployed data centers across multiple countries, including the United States, Norway, Bhutan, and Ethiopia.
To learn more, visit https://ir.bitdeer.com/
About the Role
You keep the fabric that makes 10K GPUs act like one - and turn IB/RoCE telemetry into the ground truth for our congestion and link-failure predictors.
Bitdeer is building an AI-operated GPU cloud where East-West bandwidth is the difference between a healthy training job and a $50M training run stalled by a bad optic. In this role you operate the InfiniBand and RoCEv2 fabrics that carry NCCL traffic across NeoCloud's US DCs, and you feed the AIOps substrate with the fabric telemetry it needs to catch link degradation, congestion, and topology drift before they land on the pager.
What you'll own
  • InfiniBand fabrics: fat-tree, rail-optimized, and dragonfly topologies for GPU clusters of 100-10,000 GPUs.
  • RoCEv2 networks across Nvidia, Arista, and Cisco platforms for RDMA workloads.
  • UFM (Unified Fabric Manager) for IB fabric monitoring, diagnostics, and subnet management.
  • IB and RoCE performance monitoring and tuning: adaptive routing, congestion control (DCQCN/ECN), traffic isolation.
  • NCCL communication tuning: topology detection, ring/tree algorithm selection, GDR configuration.
  • Firmware lifecycle across IB switches and HCAs.
  • Fault diagnosis: link flaps, symbol errors, packet drops, routing anomalies, credit stalls.
  • Coordination with Nvidia/Mellanox support for escalations, bugs, and RMA.

Feed the AIOps substrate
  • Wire IB/RoCE telemetry (ibdiagnet, perfquery, ibstat, PortRcvErrors, PortXmitDiscards, adaptive-routing state) into the platform's collection pipeline.
  • Partner with the platform team to define the Link and Straggler predictors: what a "bad optic 30 minutes from failure" looks like in the counters.
  • Convert every incident into a labeled example the fault-prediction engine can learn from - and every routine mitigation into a workflow the remediation actuator can run.

Job Requirement:
  • 5+ years in data center networking, with at least 3 years focused on InfiniBand or RoCE fabrics
  • Hands-on experience deploying and operating Nvidia/Mellanox InfiniBand switches at scale
  • Strong understanding of IB subnet management, partitioning, and QoS
  • Experience with RoCEv2 deployment including PFC, ECN, DCQCN configuration
  • Proficiency with UFM or equivalent IB fabric management tools
  • Knowledge of 400G/800G optics, cabling standards, and structured cabling best practices
  • Experience diagnosing IB/RoCE network issues using ibdiagnet, perfquery, ibstat, and similar tools
  • Understanding of NCCL and how GPU communication maps to network topology
  • Instinct for telemetry-driven ops - you've either built dashboards/alerts on RDMA counters at scale, or you can articulate the feature set a fabric-health model would need.
  • Runbook-as-code mindset - the diagnostics you run today should become automation next quarter.

Bitdeer is committed to providing equal employment opportunities in accordance with country, state, and local laws. Bitdeer does not discriminate against employees or applicants based on conditions such as race, color, gender identity and/or expression, sexual orientation, marital and/or parental status, religion, political opinion, nationality, ethnic background or social origin, social status, disability, age, indigenous status, and union.