... InfiniBand networking on huge training runs • Build fault tolerance systems for large-scale pretraining • Collaborate with researchers on evolving RL infrastructure • Find clean scenes in ...
... InfiniBand networking on huge training runs • Build fault tolerance systems for large-scale pretraining • Collaborate with researchers on evolving RL infrastructure • Find clean scenes in ...
Staff Hardware Systems Engineer
San Francisco, CA · On-site
$215K - $260K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Quick apply
Staff Hardware Systems Engineer
San Francisco, CA · On-site
$215K - $260K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Senior Hardware Systems Engineer
Sunnyvale, CA · On-site
$170K - $205K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Senior Hardware Systems Engineer
Sunnyvale, CA · On-site
$170K - $205K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Senior Distributed Systems Engineer
Sunnyvale, CA · On-site
$122K - $166K/yr
Required : • Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) • Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA • Deep familiarity with ...
Senior Distributed Systems Engineer
Sunnyvale, CA · On-site
$122K - $166K/yr
Required : • Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) • Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA • Deep familiarity with ...
Senior Distributed Systems Engineer
Sunnyvale, CA · On-site
$122K - $166K/yr
Required : • Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) • Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA • Deep familiarity with ...
Senior Distributed Systems Engineer
Sunnyvale, CA · On-site
$122K - $166K/yr
Required : • Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) • Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA • Deep familiarity with ...
Principal ATE Test Engineer
$209K - $240K/yr
Working knowledge of high-speed protocols like PCIe, Ethernet, Infiniband, DDR, NVMe, USB, etc. * Professional attitude with ability to execute on multiple tasks with minimal supervision. * Strong ...
Principal ATE Test Engineer
$209K - $240K/yr
Working knowledge of high-speed protocols like PCIe, Ethernet, Infiniband, DDR, NVMe, USB, etc. * Professional attitude with ability to execute on multiple tasks with minimal supervision. * Strong ...
LLM Pre-training & Distributed Engineer (AI Infrastructure)
San Francisco, CA · On-site
$126K - $166K/yr
Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors. * Automate checkpointing and failure recovery during month-long training runs. Required Skills: * Deep ...
LLM Pre-training & Distributed Engineer (AI Infrastructure)
San Francisco, CA · On-site
$126K - $166K/yr
Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors. * Automate checkpointing and failure recovery during month-long training runs. Required Skills: * Deep ...
Staff Hardware Systems Engineer
Sunnyvale, CA · On-site
$215K - $260K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Quick apply
Staff Hardware Systems Engineer
Sunnyvale, CA · On-site
$215K - $260K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Staff Hardware Systems Engineer
San Francisco, CA · On-site
$215K - $260K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Staff Hardware Systems Engineer
San Francisco, CA · On-site
$215K - $260K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Demonstrate expertise on advanced GPU and network systems (Spectrum-X, BlueField DPU, InfiniBand/RoCE, etc.) for key accounts. Run regular technical account reviews covering roadmap alignment ...
Demonstrate expertise on advanced GPU and network systems (Spectrum-X, BlueField DPU, InfiniBand/RoCE, etc.) for key accounts. Run regular technical account reviews covering roadmap alignment ...
Demonstrate expertise on advanced GPU and network systems (Spectrum-X, BlueField DPU, InfiniBand/RoCE, etc.) for key accounts. Run regular technical account reviews covering roadmap alignment ...
Demonstrate expertise on advanced GPU and network systems (Spectrum-X, BlueField DPU, InfiniBand/RoCE, etc.) for key accounts. Run regular technical account reviews covering roadmap alignment ...
Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch. * Leading proof-of-concepts and ...
Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch. * Leading proof-of-concepts and ...
Senior Hardware Systems Engineer
Sunnyvale, CA · On-site
$170K - $205K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Quick apply
Senior Hardware Systems Engineer
Sunnyvale, CA · On-site
$170K - $205K/yr
InfiniBand (fabric debugging, throughput, connectivity issues) * NVMe/storage (performance bottlenecks, firmware interactions, failure analysis) * Conduct rigorous system validation and ...
Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch. * Leading proof-of-concepts and ...
Diagnosing complex infrastructure and distributed systems issues spanning InfiniBand and RoCE fabrics, cloud interconnects, RDMA, NCCL, NVLink, and NVSwitch. * Leading proof-of-concepts and ...
LLM Pre-training & Distributed Engineer (AI Infrastructure)
San Francisco, CA · On-site
$126K - $166K/yr
Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors. * Automate checkpointing and failure recovery during month-long training runs. Required Skills: * Deep ...
LLM Pre-training & Distributed Engineer (AI Infrastructure)
San Francisco, CA · On-site
$126K - $166K/yr
Optimize networking (InfiniBand/RDMA) and memory management to prevent out-of-memory errors. * Automate checkpointing and failure recovery during month-long training runs. Required Skills: * Deep ...
Principal Software Engineer, GPU Compute
San Mateo, CA · On-site
$153K - $206K/yr
Evaluate and onboard new GPU and AI accelerator platforms, networking topologies (NVLink, InfiniBand, RoCE), and multi-node training and inference patterns. * Establish the standards, tooling, and ...
Principal Software Engineer, GPU Compute
San Mateo, CA · On-site
$153K - $206K/yr
Evaluate and onboard new GPU and AI accelerator platforms, networking topologies (NVLink, InfiniBand, RoCE), and multi-node training and inference patterns. * Establish the standards, tooling, and ...
... InfiniBand RDMA networking. • Develop containerization strategies using NVIDIA NGC, Docker, and Singularity/Apptainer. • Engineer solutions for deep learning frameworks (PyTorch, TensorFlow, JAX ...
... InfiniBand RDMA networking. • Develop containerization strategies using NVIDIA NGC, Docker, and Singularity/Apptainer. • Engineer solutions for deep learning frameworks (PyTorch, TensorFlow, JAX ...
Senior Solution Engineer, Networking
Santa Clara, CA · On-site
$65 - $83.75/hr
They provide top support for high-speed interconnect technologies like InfiniBand, NVLink, and Spectrum-X that link GPUs and AI compute infrastructure. Candidates must have a software development ...
New
Senior Solution Engineer, Networking
Santa Clara, CA · On-site
$65 - $83.75/hr
They provide top support for high-speed interconnect technologies like InfiniBand, NVLink, and Spectrum-X that link GPUs and AI compute infrastructure. Candidates must have a software development ...
New
Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...
Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...
Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...
Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...
Infiniband information
What is InfiniBand?
What are the key skills and qualifications needed to thrive as an InfiniBand network engineer, and why are they important?
What are the typical responsibilities of an InfiniBand network engineer in a data center environment?
What is the difference between Infiniband vs Ethernet Network Engineer?
| Aspect | Infiniband | Ethernet Network Engineer |
|---|---|---|
| Required Credentials | Networking certifications, Cisco, Cisco CCNA, CCNP | Networking certifications, Cisco, CCNA, CCNP |
| Work Environment | Data centers, high-performance computing environments | Corporate networks, data centers, enterprise environments |
| Industry Usage | High-performance computing, research institutions | Business, telecommunications, enterprise IT |
| Common Search/Comparison | Yes | Yes |
Infiniband and Ethernet Network Engineers both work with network infrastructure, but Infiniband specializes in high-speed, low-latency connections used in data centers and HPC environments. Ethernet Network Engineers focus on standard Ethernet networks used across various industries. While their certifications and skills overlap, their work environments and applications differ significantly.

Full-time
Re-posted 10 days ago
Job description
Krea.ai is dedicated to building next-generation AI creative tools that empower human creativity. The Engineer for Supercomputing & Distributed Systems will work on building and operating the infrastructure for research and inference, focusing on distributed training, GPU clusters, and data pipelines.
Responsibilities:
• Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets
• Run classification models on billions of images
• Deploy and combine LLMs to caption massive multimedia data
• Manage distributed training and inference on 1000+ GPU Kubernetes clusters
• Solve orchestration and scaling for large-scale GPU job processing
• Scale workloads and research between clusters in multiple datacenters
• Profile and optimize dataloaders streaming thousands of images per second
• Profile and debug InfiniBand networking on huge training runs
• Build fault tolerance systems for large-scale pretraining
• Collaborate with researchers on evolving RL infrastructure
• Find clean scenes in millions of videos using distributed shot-boundary detection
• Customize and train models to filter billions of images for questions like "is this a screenshot?"
• Build the systems that bridge raw cluster capacity and research output
Qualifications:
Required:
• Intuition for distributed systems
• Great mental model of how systems interact and function under different conditions
• Experience with Python
• Experience with Kubernetes
• Experience with Torch
• Experience with data tools like DuckDB, Arrow, etc.
Preferred:
• Experience with PyArrow
• Experience with SQL
• Experience with massive relational databases
• Experience with PyTorch
• Experience with Pandas
• Experience with NumPy
• Experience designing and implementing large-scale ETL systems
• Fundamental knowledge of containerization, operating systems, file-systems, and networking
• Experience with distributed systems design
• Experience with distributed training systems (NCCL, InfiniBand, RDMA)
• Experience with streaming and event processing systems (Kafka, Pulsar, or similar)
• Experience with PyTorch internals, custom dataloaders, and training infrastructure
Company:
Krea is a leading generative AI creative platform and research lab, offering AI tools for creatives to generate, edit, and enhance images, video, and 3D content in seconds. Founded in 2022, the company is headquartered in San Francisco, USA, with a team of 11-50 employees. The company is currently Early Stage.