RDMA/InfiniBand optimization experience * Contributions to GPU libraries or frameworks * Low-level debugging skills (PTX/SASS reading) Genmo is an Equal Opportunity Employer. Candidates are evaluated ...
Quick apply
RDMA/InfiniBand optimization experience * Contributions to GPU libraries or frameworks * Low-level debugging skills (PTX/SASS reading) Genmo is an Equal Opportunity Employer. Candidates are evaluated ...
Quick apply
RDMA/InfiniBand optimization experience * Contributions to GPU libraries or frameworks * Low-level debugging skills (PTX/SASS reading) Genmo is an Equal Opportunity Employer. Candidates are evaluated ...
Familiarity with RDMA, RoCE, InfiniBand, Ethernet, GPUDirect RDMA, or networking considerations affecting GPU cluster performance. Estimated Min Rate : $250,000.00/Annually Estimated Max Rate : $300 ...
New
Quick apply
Familiarity with RDMA, RoCE, InfiniBand, Ethernet, GPUDirect RDMA, or networking considerations affecting GPU cluster performance. Estimated Min Rate : $250,000.00/Annually Estimated Max Rate : $300 ...
New
Contribute to the design of next-generation GPU and AI interconnect fabrics, ensuring seamless ... Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design * Strong ...
Contribute to the design of next-generation GPU and AI interconnect fabrics, ensuring seamless ... Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design * Strong ...
... RDMA networking. • Develop containerization strategies using NVIDIA NGC, Docker, and Singularity ... using GPU-accelerated computing, • Develop documentation, best practices guides, and training ...
... RDMA networking. • Develop containerization strategies using NVIDIA NGC, Docker, and Singularity ... using GPU-accelerated computing, • Develop documentation, best practices guides, and training ...
Communication optimization (NCCL, RDMA, GPU interconnects) * FSDP / ZeRO and model sharding Orchestration & Runtime Systems * Ray, Kubernetes, Slurm * Distributed runtimes and async systems
Communication optimization (NCCL, RDMA, GPU interconnects) * FSDP / ZeRO and model sharding Orchestration & Runtime Systems * Ray, Kubernetes, Slurm * Distributed runtimes and async systems
San Francisco, CA · On-site
$127K - $173K/yr
By aggregating computing resources across the globe, we offer an innovative GPU marketplace and AI ... Familiarity with high-performance networking technologies such as InfiniBand and RoCE (RDMA over ...
San Francisco, CA · On-site
$127K - $173K/yr
By aggregating computing resources across the globe, we offer an innovative GPU marketplace and AI ... Familiarity with high-performance networking technologies such as InfiniBand and RoCE (RDMA over ...
San Francisco, CA · On-site
$145K - $192K/yr
When people finance GPU clusters, the datacenters housing them, and the infrastructure powering ... Knowledge of distributed training performance (NCCL, GPUDirect RDMA, multi-rail networking ...
San Francisco, CA · On-site
$145K - $192K/yr
When people finance GPU clusters, the datacenters housing them, and the infrastructure powering ... Knowledge of distributed training performance (NCCL, GPUDirect RDMA, multi-rail networking ...
Contribute to the design of next-generation GPU and AI interconnect fabrics, ensuring seamless ... Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design.
Contribute to the design of next-generation GPU and AI interconnect fabrics, ensuring seamless ... Deep familiarity with RoCEv2, RDMA transport tuning, ECN/PFC, and lossless Ethernet design.
Austin, TX · On-site
Hudson River Trading (HRT) is looking for GPU Systems Engineers to help scale and evolve our ... Experience with NVIDIA technologies beyond CUDA, such as NCCL, GPUDirect RDMA, and NVLink
Austin, TX · On-site
Hudson River Trading (HRT) is looking for GPU Systems Engineers to help scale and evolve our ... Experience with NVIDIA technologies beyond CUDA, such as NCCL, GPUDirect RDMA, and NVLink
Springfield, OH · Remote
$120K - $150K/yr
JOB TITLESenior AI GPU Deployment EngineerDEPARTMENTCloud OpsLOCATIONUSAWORK ... RDMA, RoCE, and lossless Ethernet fabricsCluster automation, observability, and lifecycle ...
Quick apply
Springfield, OH · Remote
$120K - $150K/yr
JOB TITLESenior AI GPU Deployment EngineerDEPARTMENTCloud OpsLOCATIONUSAWORK ... RDMA, RoCE, and lossless Ethernet fabricsCluster automation, observability, and lifecycle ...
$115K - $140K/yr
Advise on cluster design: multi-GPU topology, NVLink/NVSwitch considerations, RDMA, Infiniband and RoCE Ethernet, networking throughput, and storage IOPS requirements. * Guide customers in selecting ...
$115K - $140K/yr
Advise on cluster design: multi-GPU topology, NVLink/NVSwitch considerations, RDMA, Infiniband and RoCE Ethernet, networking throughput, and storage IOPS requirements. * Guide customers in selecting ...
Design and implement RDMA-based services and infrastructure that enable low-latency, high-throughput communication across GPU clusters. * Drive the evolution of collective communication frameworks ...
Design and implement RDMA-based services and infrastructure that enable low-latency, high-throughput communication across GPU clusters. * Drive the evolution of collective communication frameworks ...
The role involves conducting performance characterization and analysis on large multi-GPU and multi ... RDMA, topologies, congestion control • Experience debugging network issues in large scale ...
The role involves conducting performance characterization and analysis on large multi-GPU and multi ... RDMA, topologies, congestion control • Experience debugging network issues in large scale ...
Design and implement RDMA-based services and infrastructure that enable low-latency, high-throughput communication across GPU clusters. * Drive the evolution of collective communication frameworks ...
Design and implement RDMA-based services and infrastructure that enable low-latency, high-throughput communication across GPU clusters. * Drive the evolution of collective communication frameworks ...
Austin, TX · Hybrid
$109K - $146K/yr
The focus of this role is the RDMA networks used in AI Clusters, understanding data flows between GPU, NIC and cluster network. The ideal candidate will have a strong background in GPU architectures ...
Austin, TX · Hybrid
$109K - $146K/yr
The focus of this role is the RDMA networks used in AI Clusters, understanding data flows between GPU, NIC and cluster network. The ideal candidate will have a strong background in GPU architectures ...
Austin, TX · On-site
$142K/yr
The focus of this role is the RDMA networks used in AI Clusters, understanding data flows between GPU, NIC and cluster network. The ideal candidate will have a strong background in GPU architectures ...
Austin, TX · On-site
$142K/yr
The focus of this role is the RDMA networks used in AI Clusters, understanding data flows between GPU, NIC and cluster network. The ideal candidate will have a strong background in GPU architectures ...
Sunnyvale, CA · On-site
$122K - $166K/yr
Required : • Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) • Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA • Deep familiarity with ...
Sunnyvale, CA · On-site
$122K - $166K/yr
Required : • Experience optimizing distributed training at 1,000+ GPU scale (or equivalent depth) • Hands-on expertise with RDMA, InfiniBand, RoCE, and GPUDirect RDMA • Deep familiarity with ...
Behind that sits a large GPU fleet spread across several cloud providers. Today, our inference ... InfiniBand, RoCE, or RDMA in production. * Observability for ML workloads: Prometheus, Grafana, or ...
New
Behind that sits a large GPU fleet spread across several cloud providers. Today, our inference ... InfiniBand, RoCE, or RDMA in production. * Observability for ML workloads: Prometheus, Grafana, or ...
New
Knowledge of GPUDirect RDMA and GPU-aware communication technologies. * Experience developing congestion management, traffic engineering, or network resiliency solutions. * Familiarity with large ...
Knowledge of GPUDirect RDMA and GPU-aware communication technologies. * Experience developing congestion management, traffic engineering, or network resiliency solutions. * Familiarity with large ...
New York, NY · On-site
$165K - $330K/yr
Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack, helping us move beyond TCP/IP to unlock order-of-magnitude improvements in ...
New York, NY · On-site
$165K - $330K/yr
Make RDMA First-Class: You will work on integrating RDMA/RoCE/InfiniBand capabilities directly into our inference stack, helping us move beyond TCP/IP to unlock order-of-magnitude improvements in ...
$42.5K - $54.5K
0% of jobs
$54.5K - $66.6K
0% of jobs
$66.6K - $78.6K
2% of jobs
$78.6K - $90.7K
7% of jobs
$90.7K - $102.7K
13% of jobs
$104.9K is the 25th percentile. Wages below this are outliers.
$102.7K - $114.8K
18% of jobs
The median wage is $121.5K / yr.
$114.8K - $126.8K
19% of jobs
$126.8K - $138.9K
16% of jobs
$140K is the 75th percentile. Wages above this are outliers.
$138.9K - $150.9K
11% of jobs
$150.9K - $163K
9% of jobs
$163K - $175K
5% of jobs
$42.5K
$123.8K
$175K
| Aspect | Rdma Gpu | Network Engineer |
|---|---|---|
| Required Credentials | Computer science or related degree, certifications in GPU computing or high-performance networking | Networking certifications (CCNA, CCNP), degree in computer science or related field |
| Work Environment | Data centers, high-performance computing labs, research facilities | Corporate offices, data centers, telecommunication environments |
| Industry Usage | AI, machine learning, scientific computing, data analytics | IT infrastructure, network design, security, and maintenance |
Rdma Gpu specialists focus on optimizing GPU performance and high-speed data transfer using RDMA technology, primarily in computing and research environments. Network Engineers design, implement, and maintain network systems. While both roles involve high-tech infrastructure, Rdma Gpu roles are more specialized in GPU and high-performance data transfer, whereas Network Engineers focus on network connectivity and security.
Cities with the most Rdma Gpu job openings:
States with the most job openings for Rdma Gpu jobs include:
The top searched job categories for Rdma Gpu jobs are:

Full-time
Re-posted 11 days ago
We are Genmo, a research lab dedicated to building open, state-of-the-art models for video generation towards unlocking the right brain of AGI. Join us in shaping the future of AI and pushing the boundaries of what's possible in video generation.
We're seeking a GPU Performance Engineer to squeeze every last FLOP from our H100 infrastructure and optimize our model serving stack to its absolute limits.
The Role
You'll be our performance optimization expert, using advanced profiling tools to identify bottlenecks and implementing solutions that achieve 5-10x speedups. From writing custom CUDA kernels to eliminating cold start latency, you'll ensure our infrastructure delivers world-class performance. This role is perfect for someone who gets excited about microsecond optimizations and pushing hardware to its theoretical limits.
Key Responsibilities
Profile and optimize GPU workloads using Nsight Systems, nvprof, and custom instrumentation
Write high-performance CUDA and Triton kernels for critical model operations
Optimize cold start latency from seconds to milliseconds for our serving infrastructure
Tune memory access patterns, kernel fusion, and GPU utilization
Collaborate with ML engineers to optimize model implementations
Debug performance issues across the full stack from application to hardware
Implement custom memory pooling and allocation strategies
Share optimization techniques and build performance culture across teams
Qualifications
Bachelor's or Master's degree in Computer Science, Electrical Engineering, or related field
5+ years systems programming experience with 3+ years focused on GPU optimization
Expert proficiency with GPU profiling tools (Nsight Systems, nvprof)
Strong CUDA programming skills with production kernel development
Deep understanding of GPU architecture (memory hierarchy, SMs, warps)
Track record of achieving significant performance improvements (5-10x)
Experience with Python and C++ in production environments
We Value
Experience with Triton kernel development
Knowledge of CUTLASS or similar high-performance libraries
Background in ML-specific optimizations (attention, transformers)
RDMA/InfiniBand optimization experience
Contributions to GPU libraries or frameworks
Low-level debugging skills (PTX/SASS reading)
Genmo is an Equal Opportunity Employer. Candidates are evaluated without regard to age, race, color, religion, sex, disability, national origin, sexual orientation, veteran status, or any other characteristic protected by federal or state law. Genmo, Inc. is an E-Verify company and you may review the Notice of E-Verify Participation and the Right to Work posters in English and Spanish.