1

Infiniband Jobs in Michigan (NOW HIRING)

InfiniBand / high-speed networking * PXE provisioning and configuration management * PHP/web-based tools or dashboards * HPC performance tuning and capacity planning Ideal Candidate: Senior Linux/HPC ...

New

Infiniband information

What is InfiniBand?

Infiniband is a high-speed, low-latency networking technology commonly used in data centers and high-performance computing environments. It is designed to connect servers, storage systems, and network devices, providing much faster data transfer rates than traditional Ethernet. Infiniband supports scalable bandwidth and efficient communication, which makes it ideal for applications requiring rapid data movement, such as scientific simulations and large-scale database transactions. Its architecture also supports remote direct memory access (RDMA), which further reduces latency and CPU overhead.

What are the typical responsibilities of an InfiniBand network engineer in a data center environment?

InfiniBand network engineers are primarily responsible for designing, deploying, and maintaining high-performance InfiniBand fabrics that connect servers and storage systems in data centers, especially in HPC (High-Performance Computing) environments. Their daily tasks include monitoring network performance, troubleshooting connectivity or latency issues, and performing firmware and driver updates on InfiniBand switches and host adapters. They also collaborate closely with system administrators and application teams to optimize throughput and ensure reliable, low-latency communication. Additionally, InfiniBand engineers often participate in capacity planning and help scale the network infrastructure to meet growing computational demands.

What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?

To thrive as an InfiniBand Network Engineer, you need a strong background in computer networking, Linux system administration, and high-performance computing (HPC) environments, often supported by a degree in computer science or related field. Familiarity with InfiniBand architecture, experience with tools like OpenFabrics Enterprise Distribution (OFED), and certifications such as CompTIA Network+ are valuable. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for this role. These abilities are essential for ensuring efficient, reliable InfiniBand network performance in complex HPC or data center environments.

What is the difference between Infiniband vs Ethernet Network Engineer?

AspectInfinibandEthernet Network Engineer
Required CredentialsNetworking certifications, Cisco, Cisco CCNA, CCNPNetworking certifications, Cisco, CCNA, CCNP
Work EnvironmentData centers, high-performance computing environmentsCorporate networks, data centers, enterprise environments
Industry UsageHigh-performance computing, research institutionsBusiness, telecommunications, enterprise IT
Common Search/ComparisonYesYes

Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

What cities in Michigan are hiring for Infiniband jobs?

Cities in Michigan with the most Infiniband job openings:

Infographic showing various Infiniband job openings in Michigan as of August 2026, with employment types broken down into 92% Full Time, 3% Part Time, and 5% Contract. Highlights an 85% Physical, 5% Hybrid, and 10% Remote job distribution.

CAE HPC System Administrator

Saline, MI

Contractor

Re-posted 3 days ago


Job description

About TTS-US:

Founded in 2011, Toyota Tsusho Systems US, Inc. (TTS-US) is a Toyota group company, that develops IT solutions wherever global businesses operate. Transforming into a technology and mobility company, TTS-US, with its 8 TTS affiliates worldwide is establishing a secure and resilient Toyota global value chain. The creative capacity to forge such limitless business opportunities is one of the strengths of Toyota Tsusho Systems."                                                                                                                          

Position Summary:                                                                                                               

We are seeking a highly motivated and experienced CAE HPC System Administrator with more than 7 years of experience to join our dynamic digital solution manufacturing team. This position is ideal for a candidate with strong Linux system administration experience and hands-on expertise managing HPC environments for CAE workloads, including job schedulers, system automation, and engineering application support. The successful candidate will be responsible for administering and optimizing HPC clusters, managing job scheduling systems, supporting CAE applications and licensing, automating Linux operations, maintaining infrastructure performance, and ensuring system stability, scalability, and efficient workload execution.                                                                                 

Requirements

Essential Functions:                                                                                                                         

  1. HPC Job Queuing & Workload Management
    • Administer, configure, and optimize HPC job scheduling environments, including IBM Spectrum LSF, Open PBS,  or equivalent schedulers.
    • Design and tune job queues, resource allocation policies, and scheduling strategies to support diverse CAE workloads.
    • Monitor system performance and utilization trends and implement improvements to maximize efficiency and throughput.
  2. CAE Application and Licensing Support
    • Install, upgrade, test, and support CAE applications and simulation tools in production environments.
    • Provide integration support between CAE applications and HPC scheduling systems.
    • Manage CAE software licensing systems (e.g., FlexLM, RLM) and ensure availability.
    • Troubleshoot application-related issues and ensure minimal disruption to engineering activities.
  3. Linux Systems Administration & Automation
    • Administer and maintain Red Hat Enterprise Linux (RHEL) environments across HPC clusters.
    • Perform OS provisioning, deployment, and patch management using automated tools (e.g., PXE, or configuration management solutions).
    • Develop and maintain scripts (Bash, Korn shell, C Shell, Perl, Awk, or equivalent) to automate system monitoring, health checks, and routine administrative tasks.
    • Maintain system logs, monitoring processes, and standard operating procedures.
  4. Hardware & Infrastructure Management
    • Troubleshoot and resolve issues related to servers, storage systems, and high-performance networking (e.g., InfiniBand, high-speed Ethernet).
    • Support hardware lifecycle activities including installation, maintenance, and upgrades.
    • Conduct capacity planning based on system utilization trends and future demand.
  5. Operations, Monitoring & Continuous Improvement
    • Perform system health checks, monitoring, and incident tracking for HPC and CAE environments.
    • Document system configurations, procedures, incidents, and best practices.
    • Track outages, analyze root causes, and implement preventive measures.
    • Follow change management processes for system updates and deployments.
    • Provide accurate reporting (e.g., utilization, incidents, system performance) and support project initiatives.

Minimum qualifications:                                                                                                    

Required Education & Experience:

7+ years of Linux system administration experience (preferably RHEL environments).

Bachelor's degree in mechanical engineering, electrical engineering, computer engineering, computer science, or related field; and/or commensurate work experience

Hands-on experience managing HPC clusters and job schedulers (LSF, Slurm, PBS, or similar).

Proven experience in CAE application support and integration.

Strong scripting skills (Bash, Shell, Perl, or equivalent).

Experience with OS deployment, patching, and system automation.

Solid understanding of enterprise server hardware, storage, and networking fundamentals.

Experience with CAE tools such as Ansys, LS-DYNA, Nastran, or similar.

Familiarity with high-performance networking technologies is plus (e.g., InfiniBand).

Experience developing internal tools or dashboards are plus (e.g., PHP or web-based tooling).

Position Type/Expected Hours of Work:

Hybrid Full-time contract: Standard business hours with flexibility required to support maintenance windows and critical production issues.

Occasional after-hours or weekend work may be required based on business needs.