... UCX, NVSHMEM) • Experience conducting performance benchmarking and triage on large scale HPC clusters • Good understanding of computer system architecture, HW-SW interactions and operating ...

60 Ucx Jobs Hiring Near You
... UCX, NVSHMEM) • Experience conducting performance benchmarking and triage on large scale HPC clusters • Good understanding of computer system architecture, HW-SW interactions and operating ...
Solutions Architect, Inference Deployments
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Solutions Architect, Inference Deployments
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Senior Distributed Systems Engineer
Sunnyvale, CA · On-site
$122K - $166K/yr
... or UCX internals • Strong systems programming ability (C/C++, Rust, or Go) • Strong familiarity with modern model training frameworks such as PyTorch • Ability to troubleshoot and profile ...
Senior Distributed Systems Engineer
Sunnyvale, CA · On-site
$122K - $166K/yr
... or UCX internals • Strong systems programming ability (C/C++, Rust, or Go) • Strong familiarity with modern model training frameworks such as PyTorch • Ability to troubleshoot and profile ...
Solutions Architect, Inference Deployments
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Solutions Architect, Inference Deployments
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Solutions Architect, Inference Deployments
Santa Clara, CA · On-site
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Solutions Architect, Inference Deployments
Santa Clara, CA · On-site
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Senior Deep Learning Communication Architect
Austin, TX · On-site
$128K - $174K/yr
Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...
Senior Deep Learning Communication Architect
Austin, TX · On-site
$128K - $174K/yr
Collaborate with hardware and software teams to craft systems that effectively apply high-speed interconnects (e.g., NVLink, InfiniBand, SPC-X) and communication libraries (e.g., MPI, NCCL, UCX, UCC ...
Solutions Architect, Inference Deployments
Santa Clara, CA · On-site
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Solutions Architect, Inference Deployments
Santa Clara, CA · On-site
$74 - $97.50/hr
Solving sophisticated GPU allocation, memory hierarchies (HBM, DRAM, SSD), and low-latency networking (RDMA, UCX). * Demonstrated success in tuning large language models for low-latency inference in ...
Senior Software Engineer, CUDA Deep Learning Systems
Austin, TX · On-site
$121K - $160K/yr
... UCX) and distributed machine learning techniques (e.g., pipeline parallelism, tensor parallelism). • Knowledge of numerical methods, low-precision arithmetic (e.g., NVFP4, MXFP4, FP8, INT8), and ...
Senior Software Engineer, CUDA Deep Learning Systems
Austin, TX · On-site
$121K - $160K/yr
... UCX) and distributed machine learning techniques (e.g., pipeline parallelism, tensor parallelism). • Knowledge of numerical methods, low-precision arithmetic (e.g., NVFP4, MXFP4, FP8, INT8), and ...
Implementing lower-level communication frameworks like UCX and libfabric, or development using RDMA APIs * Development and optimization of communication collective algorithms (e.g. AllReduce)
Implementing lower-level communication frameworks like UCX and libfabric, or development using RDMA APIs * Development and optimization of communication collective algorithms (e.g. AllReduce)
Senior System Software Engineer
Santa Clara, CA · On-site
$143K - $189K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
Senior System Software Engineer
Santa Clara, CA · On-site
$143K - $189K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
Principal Systems Software Engineer
Santa Clara, CA · On-site
$158K - $212K/yr
Work with open source communities to enhance libraries like NVIDIA cuDF, CCCL and UCX through technical discussion and code contributions * Collaborate with distributed systems teams to craft ...
Principal Systems Software Engineer
Santa Clara, CA · On-site
$158K - $212K/yr
Work with open source communities to enhance libraries like NVIDIA cuDF, CCCL and UCX through technical discussion and code contributions * Collaborate with distributed systems teams to craft ...
Senior System Software Engineer
$137K - $180K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
Senior System Software Engineer
$137K - $180K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
Senior Distributed Systems Engineer
Sunnyvale, CA · Hybrid
$200K - $400K/yr
... UCX internals · Strong systems programming ability (C/C++, Rust, or Go) · Strong familiarity with modern model training frameworks such as PyTorch · Ability to troubleshoot and profile training ...
Quick apply
Senior Distributed Systems Engineer
Sunnyvale, CA · Hybrid
$200K - $400K/yr
... UCX internals · Strong systems programming ability (C/C++, Rust, or Go) · Strong familiarity with modern model training frameworks such as PyTorch · Ability to troubleshoot and profile training ...
Principal Systems Software Engineer
Champaign, IL · On-site
$135K - $181K/yr
Work with open source communities to enhance libraries like NVIDIA cuDF, CCCL and UCX through technical discussion and code contributions * Collaborate with distributed systems teams to craft ...
Principal Systems Software Engineer
Champaign, IL · On-site
$135K - $181K/yr
Work with open source communities to enhance libraries like NVIDIA cuDF, CCCL and UCX through technical discussion and code contributions * Collaborate with distributed systems teams to craft ...
Principal Systems Software Engineer
Santa Clara, CA · On-site
$158K - $212K/yr
Work with open source communities to enhance libraries like NVIDIA cuDF, CCCL and UCX through technical discussion and code contributions * Collaborate with distributed systems teams to craft ...
Principal Systems Software Engineer
Santa Clara, CA · On-site
$158K - $212K/yr
Work with open source communities to enhance libraries like NVIDIA cuDF, CCCL and UCX through technical discussion and code contributions * Collaborate with distributed systems teams to craft ...
Senior System Software Engineer
Santa Clara, CA · On-site
$143K - $189K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
Senior System Software Engineer
Santa Clara, CA · On-site
$143K - $189K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
Senior Software Engineer, NCCL
$143K - $189K/yr
UCX for MPI/OpenSHMEM) on GPU clusters. * Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. * Design, implement and maintain system software that ...
Senior Software Engineer, NCCL
$143K - $189K/yr
UCX for MPI/OpenSHMEM) on GPU clusters. * Participating in and contributing to parallel programming interface specifications like MPI/OpenSHMEM. * Design, implement and maintain system software that ...
Senior System Software Engineer
$122K - $161K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
Senior System Software Engineer
$122K - $161K/yr
Work with open source communities to enhance libraries like RAPIDS, CCCL and UCX through technical discussion and code contributions * Provide recommendations and feedback to teams regarding ...
NIXL, NCCL, UCX, MPI, NVSHMEM), and GPU accelerated systems, with track record of defining and delivering complex, cross-team technical initiatives from research concept to production. * MS, PhD or ...
NIXL, NCCL, UCX, MPI, NVSHMEM), and GPU accelerated systems, with track record of defining and delivering complex, cross-team technical initiatives from research concept to production. * MS, PhD or ...
NIXL, NCCL, UCX, MPI, NVSHMEM), and GPU accelerated systems, with track record of defining and delivering complex, cross-team technical initiatives from research concept to production. * MS, PhD or ...
NIXL, NCCL, UCX, MPI, NVSHMEM), and GPU accelerated systems, with track record of defining and delivering complex, cross-team technical initiatives from research concept to production. * MS, PhD or ...
UCX Jobs Information

Full-time
Posted 5 days ago
Nvidia rating
9.6
Based on 17 frontline employees who took The Breakroom Quiz
5th of 217 rated software companies
Job description
NVIDIA is leading the way in groundbreaking developments in Artificial Intelligence, High Performance Computing and Visualization. The role involves conducting performance characterization and analysis on large multi-GPU and multi-node clusters to influence the roadmap of communication libraries.
Responsibilities:
• Conduct in-depth performance characterization and analysis on large multi-GPU and multi-node clusters.
• Study the interaction of our libraries with all HW (GPU, CPU, Networking) and SW components in the stack
• Evaluate proof-of-concepts, conduct trade-off analysis when multiple solutions are available
• Triage and root-cause performance issues reported by our customers
• Collect a lot of performance data; build tools and infrastructure to visualize and analyze the information
• Collaborate with a very dynamic team across multiple time zones
Qualifications:
Required:
• M.S. (or equivalent experience) or PhD in Computer Science, or related field with relevant performance engineering and HPC experience
• 3+ yrs of experience with parallel programming and at least one communication runtime (MPI, NCCL, UCX, NVSHMEM)
• Experience conducting performance benchmarking and triage on large scale HPC clusters
• Good understanding of computer system architecture, HW-SW interactions and operating systems principles (aka systems software fundamentals)
• Implement micro-benchmarks in C/C++, read and modify the code base when required
• Ability to debug performance issues across the entire HW/SW stack. Proficient in a scripting language, preferably Python
• Familiar with containers, cloud provisioning and scheduling tools (Kubernetes, SLURM, Ansible, Docker)
• Adaptability and passion to learn new areas and tools. Flexibility to work and communicate effectively across different teams and timezones
Preferred:
• Practical experience with Infiniband/Ethernet networks in areas like RDMA, topologies, congestion control
• Experience debugging network issues in large scale deployments
• Familiarity with CUDA programming and/or GPUs
• Experience with Deep Learning Frameworks such PyTorch, TensorFlow
Company:
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.
About Nvidia
Sourced by ZipRecruiter
NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.
Industry
Computer and electronic product manufacturing
Company size
10,000+ Employees
Headquarters location
Santa Clara, CA, US
Year founded
1993