1

Rdma Gpu Jobs (NOW HIRING)

... as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency. • Lead efforts to improve the reliability, durability ...

GPU Fabric Engineer

$125K - $135K/yr

Optimize RDMA transport settings, PFC/ECN behavior, and lossless queue configuration for GPU traffic * Validate fabric performance benchmarks and ensure line-rate throughput for AI workloads

Collaborate with other teams to architect RDMA-capable hardware and define transport layer optimizations for GPU-based large scale AI workload deployments. * Use and modify system models, perform ...

... as RDMA, GPU Direct Storage, and distributed filesystems protocols such as NFS or FUSE to optimize storage performance and efficiency. • Lead efforts to improve the reliability, durability ...

Software Engineer, GPU Infrastructure (HPC)

$110K - $144K/yr

... like RDMA, NCCL, and high-speed interconnects. • Troubleshoot and resolve complex issues ... Experience with GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and ...

RDMA/InfiniBand optimization experience * Contributions to GPU libraries or frameworks * Low-level debugging skills (PTX/SASS reading) Genmo is an Equal Opportunity Employer. Candidates are evaluated ...

RDMA/InfiniBand optimization experience * Contributions to GPU libraries or frameworks * Low-level debugging skills (PTX/SASS reading) Genmo is an Equal Opportunity Employer. Candidates are evaluated ...

Infrastructure Engineer (Storage)

$110K - $144K/yr

... g., RDMA, GPU Direct Storage) • Experience supporting GPU-based workloads or large-scale compute clusters Company : The AI development platform - From idea to AI, Lightning fast ⚡️. Code ...

next page

Showing results 1-20

Rdma Gpu information

See salary details

$42.5K

$123.8K

$175K

How much do rdma gpu jobs pay per year?

As of Aug 6, 2026, the average yearly pay for rdma gpu in the United States is $123,786.00, according to ZipRecruiter salary data. Most workers in this role earn between $104,000.00 and $142,500.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as an RDMA GPU engineer, and why are they important?

To thrive as an RDMA GPU Engineer, you need a solid background in computer science or engineering, with expertise in GPU architectures, networking protocols, and RDMA (Remote Direct Memory Access) technologies. Familiarity with CUDA, InfiniBand, RoCE, and relevant profiling or debugging tools is typically required. Strong problem-solving, teamwork, and communication skills help in collaborating effectively on complex, performance-critical systems. These skills are crucial for optimizing high-performance computing applications and ensuring efficient data transfer between GPUs and networked devices.

What is an RDMA GPU?

RDMA GPUs are graphics processing units that support Remote Direct Memory Access (RDMA) technology, enabling direct memory transfers between the GPU and remote devices or other GPUs across a network without involving the host CPU. This technology is commonly used in high-performance computing, AI, and data centers to reduce latency and increase data throughput. RDMA GPUs allow faster data exchange for distributed computing tasks, such as large-scale machine learning training, by bypassing traditional network bottlenecks.

How does an RDMA GPU engineer typically collaborate with software and hardware teams to optimize performance?

As an RDMA GPU engineer, you’ll regularly work alongside both software developers and hardware architects to ensure that high-speed data transfers between GPUs and other components are efficient and reliable. This collaboration often involves troubleshooting bottlenecks, tuning drivers, and optimizing memory access patterns. You may also participate in code reviews, joint debugging sessions, and performance benchmarking to align system-level improvements. Effective communication across multidisciplinary teams is essential to deliver best-in-class solutions for demanding workloads like machine learning or scientific computing.

What is the difference between Rdma Gpu vs Network Engineer?

AspectRdma GpuNetwork Engineer
Required CredentialsComputer science or related degree, certifications in GPU computing or high-performance networkingNetworking certifications (CCNA, CCNP), degree in computer science or related field
Work EnvironmentData centers, high-performance computing labs, research facilitiesCorporate offices, data centers, telecommunication environments
Industry UsageAI, machine learning, scientific computing, data analyticsIT infrastructure, network design, security, and maintenance

Rdma Gpu specialists focus on optimizing GPU performance and high-speed data transfer using RDMA technology, primarily in computing and research environments. Network Engineers design, implement, and maintain network systems. While both roles involve high-tech infrastructure, Rdma Gpu roles are more specialized in GPU and high-performance data transfer, whereas Network Engineers focus on network connectivity and security.

More about Rdma Gpu jobs
What cities are hiring for Rdma Gpu jobs? Cities with the most Rdma Gpu job openings:
What states have the most Rdma Gpu jobs? States with the most job openings for Rdma Gpu jobs include:
Infographic showing various Rdma Gpu job openings in the United States as of August 2026, with employment types broken down into 95% Full Time, and 5% Contract. Highlights an 53% In-person, 5% Hybrid, and 42% Remote job distribution, with an average salary of $123,786 per year, or $59.5 per hour.

Senior GPU Inference Performance Engineer

Advanced Micro Devices, Inc

Santa Clara, CA • On-site

$123K - $169K/yr

Full-time

Posted 27 days ago


Advanced Micro Devices rating

8.6

Company rating: 8.6 out of 10

Based on 13 frontline employees who took The Breakroom Quiz

27th of 156 rated electronics manufacturers


Job description


WHAT YOU DO AT AMD CHANGES EVERYTHING 

At AMD, our mission is to build great products that accelerate next-generation computing experiences—from AI and data centers, to PCs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join AMD, you’ll discover the real differentiator is our culture. We push the limits of innovation to solve the world’s most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of AI and beyond.  Together, we advance your career.  



THE ROLE:
We are looking for a Senior GPU Inference Performance Engineer to own end-to-end performance analysis of GPU-accelerated AI inference workloads. You will profile, diagnose, and explain performance across the full stack from GPU silicon, communication libraries, networking fabrics, and operating systems through the software runtime and drive competitive positioning against other accelerator vendors. This role sits at the intersection of hardware, systems software, networking, and AI infrastructure, and requires someone who can go deep on a trace and present findings to product and executive stakeholders.
 
THE PERSON:
A hands-on performance engineer who is equally comfortable reading a GPU trace, debugging distributed systems performance issues, and briefing executives. You are curious, evidence-driven, rigorous, and you don't stop at "X is faster" and you explain why, rooted in hardware and software evidence. You collaborate across hardware, systems software, networking, and AI infrastructure teams, communicate clearly in written reports and presentations, and thrive at the intersection of silicon, operating systems, communication libraries, networking, and AI. Experience with Linux systems, distributed GPU infrastructure, RDMA/RoCE networking, or communication libraries such as NCCL/RCCL is highly valued.
 
KEY RESPONSIBILITIES:
  • Full-stack GPU profiling: Instrument and analyze inference workloads across AMD Instinct (ROCm, rocProfiler, ROCm Systems Profiler, RGP, rocprof-compute, rocprof-sys, Omniperf) and NVIDIA (CUDA, Nsight Systems/Compute, DCGM) GPUs. Identify bottlenecks spanning HBM bandwidth, compute utilization, kernel scheduling, memory allocation, PCIe/Infinity Fabric data movement, and GPU runtime behavior.
  • Systems and runtime performance analysis: Profile and diagnose performance interactions between GPU runtimes, Linux operating systems, device drivers, container runtimes, memory subsystems, CPU scheduling, NUMA topology, and I/O pathways. Identify system-level bottlenecks that impact throughput, latency, and GPU utilization.
  • Competitive performance analysis: Design and execute head-to-head benchmarks (AMD vs. NVIDIA) on standardized AI and LLM workloads. Produce clear, data-backed explanations of why performance differs attributing gaps to hardware architecture, networking topology, communication libraries, software maturity, runtime behavior, or configuration differences.
  • Multi-server inference networking: Profile and optimize distributed inference topologies including prefill-decode (PD) disaggregation, pipeline parallelism, and tensor parallelism across multi-node clusters. Analyze network-level bottlenecks using RDMA/RoCE traces, NCCL/RCCL collective profiling, GPUDirect RDMA, NIC-level counters (Pensando, ConnectX), and network performance tools. Quantify the impact of latency, bandwidth, congestion, and topology on end-to-end inference SLAs.
  • GPU operator and Kubernetes stack: Profile the overhead introduced by GPU operators, device plugins, container runtimes (Docker, containerd), and Kubernetes scheduling on inference latency. Identify and resolve jitter, cold-start, resource contention, and infrastructure inefficiencies in production environments.
  • Tooling and automation: Build reproducible benchmarking harnesses, profiling scripts, and performance regression dashboards. Automate trace collection and analysis to support continuous performance validation across firmware, drivers, networking stacks, runtimes, and AI frameworks.
PREFERRED EXPERIENCE:
  • Background in GPU performance engineering, HPC, distributed systems, networking, operating systems, or systems performance analysis.
  • Hands-on proficiency with either AMD (ROCm, rocProfiler, ROCm Systems Profiler, RGP, rocprof-compute, rocprof-sys, Omniperf/Omnitrace) or NVIDIA (CUDA, Nsight Systems/Compute, NCU) profiling toolchains, with deep understanding of GPU architecture: warp/wavefront execution, memory hierarchy, occupancy, and instruction-level parallelism.
  • Experience analyzing GPU communication and networking performance including NCCL/RCCL, RDMA/RoCE, GPUDirect RDMA, UCX, MPI, ConnectX, Pensando, or similar high-performance networking technologies.
  • Experience with multi-GPU and multi-node inference, training, or HPC environments including tensor parallelism, pipeline parallelism, distributed communication libraries, and network performance analysis tools.
  • Experience with Linux systems performance analysis, operating systems, device drivers, virtualization, container runtimes, or low-level systems software development.
  • Demonstrated ability to explain performance differences in written reports or presentations—not just "X is faster" but why, rooted in hardware and software evidence.
  • Strong Python and C/C++ skills; comfort reading GPU kernel code (HIP/CUDA), runtime code, or systems-level software.
  • Experience with Kubernetes GPU scheduling, MIG, GPU operator performance, or contributions to open-source infrastructure, systems, networking, inference, or profiling projects.
ACADEMIC CREDENTIALS:
  • Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field preferred; advanced degree desired.

This role is not eligible for visa sponsorship.

#LI-TB1

#LI-Hybrid



Benefits offered are described:  AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD’s “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.

Qualifications:

Benefits offered are described:  AMD benefits at a glance.

AMD does not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services. AMD and its subsidiaries are equal opportunity, inclusive employers and will consider all applicants without regard to age, ancestry, color, marital status, medical condition, mental or physical disability, national origin, race, religion, political and/or third-party affiliation, sex, pregnancy, sexual orientation, gender identity, military or veteran status, or any other characteristic protected by law.   We encourage applications from all qualified candidates and will accommodate applicants’ needs under the respective laws throughout all stages of the recruitment and selection process.

AMD may use Artificial Intelligence to help screen, assess or select applicants for this position.  AMD’s “Responsible AI Policy” is available here.

 

This posting is for an existing vacancy.

Education:UNAVAILABLEEmployment Type: FULL_TIME

What Advanced Micro Devices employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom