1

Infiniband Jobs in California (NOW HIRING)

InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...

InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...

Working knowledge of high-speed protocols like PCIe, Ethernet, Infiniband, DDR, NVMe, USB, etc. * Professional attitude with ability to execute on multiple tasks with minimal supervision. * Strong ...

InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...

InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...

InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...

Principal Software Engineer, GPU Compute

San Mateo, CA · On-site

$153K - $206K/yr

Evaluate and onboard new GPU and AI accelerator platforms, networking topologies (NVLink, InfiniBand, RoCE), and multi-node training and inference patterns. * Establish the standards, tooling, and ...

Senior Solution Engineer, Networking

Santa Clara, CA · On-site

$65 - $83.75/hr

They provide top support for high-speed interconnect technologies like InfiniBand, NVLink, and Spectrum-X that link GPUs and AI compute infrastructure. Candidates must have a software development ...

New

Showing results 21-40

Infiniband information

What is InfiniBand?

Infiniband is a high-speed, low-latency networking technology commonly used in data centers and high-performance computing environments. It is designed to connect servers, storage systems, and network devices, providing much faster data transfer rates than traditional Ethernet. Infiniband supports scalable bandwidth and efficient communication, which makes it ideal for applications requiring rapid data movement, such as scientific simulations and large-scale database transactions. Its architecture also supports remote direct memory access (RDMA), which further reduces latency and CPU overhead.

What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?

To thrive as an InfiniBand Network Engineer, you need a strong background in computer networking, Linux system administration, and high-performance computing (HPC) environments, often supported by a degree in computer science or related field. Familiarity with InfiniBand architecture, experience with tools like OpenFabrics Enterprise Distribution (OFED), and certifications such as CompTIA Network+ are valuable. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for this role. These abilities are essential for ensuring efficient, reliable InfiniBand network performance in complex HPC or data center environments.

What are the typical responsibilities of an InfiniBand network engineer in a data center environment?

InfiniBand network engineers are primarily responsible for designing, deploying, and maintaining high-performance InfiniBand fabrics that connect servers and storage systems in data centers, especially in HPC (High-Performance Computing) environments. Their daily tasks include monitoring network performance, troubleshooting connectivity or latency issues, and performing firmware and driver updates on InfiniBand switches and host adapters. They also collaborate closely with system administrators and application teams to optimize throughput and ensure reliable, low-latency communication. Additionally, InfiniBand engineers often participate in capacity planning and help scale the network infrastructure to meet growing computational demands.

What is the difference between Infiniband vs Ethernet Network Engineer?

AspectInfinibandEthernet Network Engineer
Required CredentialsNetworking certifications, Cisco, Cisco CCNA, CCNPNetworking certifications, Cisco, CCNA, CCNP
Work EnvironmentData centers, high-performance computing environmentsCorporate networks, data centers, enterprise environments
Industry UsageHigh-performance computing, research institutionsBusiness, telecommunications, enterprise IT
Common Search/ComparisonYesYes

Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

What job categories do people searching Infiniband jobs in California look for? The top searched job categories for Infiniband jobs in California are:
What cities in California are hiring for Infiniband jobs? Cities in California with the most Infiniband job openings:
Infographic showing various Infiniband job openings in California as of August 2026, with employment types broken down into 87% Full Time, 11% Part Time, and 2% Contract. Highlights an 86% Physical, 3% Hybrid, and 11% Remote job distribution.

Engineer, Supercomputing & Distributed Systems

krea.ai

San Francisco, CA • On-site

Full-time

Re-posted 10 days ago


Job description

Job Summary:
Krea.ai is dedicated to building next-generation AI creative tools that empower human creativity. The Engineer for Supercomputing & Distributed Systems will work on building and operating the infrastructure for research and inference, focusing on distributed training, GPU clusters, and data pipelines.
Responsibilities:
• Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets
• Run classification models on billions of images
• Deploy and combine LLMs to caption massive multimedia data
• Manage distributed training and inference on 1000+ GPU Kubernetes clusters
• Solve orchestration and scaling for large-scale GPU job processing
• Scale workloads and research between clusters in multiple datacenters
• Profile and optimize dataloaders streaming thousands of images per second
• Profile and debug InfiniBand networking on huge training runs
• Build fault tolerance systems for large-scale pretraining
• Collaborate with researchers on evolving RL infrastructure
• Find clean scenes in millions of videos using distributed shot-boundary detection
• Customize and train models to filter billions of images for questions like "is this a screenshot?"
• Build the systems that bridge raw cluster capacity and research output
Qualifications:
Required:
• Intuition for distributed systems
• Great mental model of how systems interact and function under different conditions
• Experience with Python
• Experience with Kubernetes
• Experience with Torch
• Experience with data tools like DuckDB, Arrow, etc.
Preferred:
• Experience with PyArrow
• Experience with SQL
• Experience with massive relational databases
• Experience with PyTorch
• Experience with Pandas
• Experience with NumPy
• Experience designing and implementing large-scale ETL systems
• Fundamental knowledge of containerization, operating systems, file-systems, and networking
• Experience with distributed systems design
• Experience with distributed training systems (NCCL, InfiniBand, RDMA)
• Experience with streaming and event processing systems (Kafka, Pulsar, or similar)
• Experience with PyTorch internals, custom dataloaders, and training infrastructure
Company:
Krea is a leading generative AI creative platform and research lab, offering AI tools for creatives to generate, edit, and enhance images, video, and 3D content in seconds. Founded in 2022, the company is headquartered in San Francisco, USA, with a team of 11-50 employees. The company is currently Early Stage.