2

Remote Nvidia Engineering Jobs in California (NOW HIRING)

Trusted by organizations including NVIDIA, HPE, the London Stock Exchange, PSI CRO, the U.S. Air ... Location Remote - United States preferably Bay Area location About the Role Step into Arango ...

Software Engineer, Data Engineering

San Francisco, CA · On-site +1

$134K - $162K/yr

United States (remote) What is Verse? The race to AI has become the race to power. Every ... Backed by Bessemer Venture Partners, GV, Coatue, and NVIDIA, and built by pioneers in grid-scale ...

Senior Software Engineer - CUDA

Palo Alto, CA · On-site +1

$144K - $189K/yr

In this role, you will collaborate with our engineering team to identify performance bottlenecks ... Experience with profiling and debugging tools for GPU applications, such as NVIDIA Nsight.

Senior Software Engineer - CUDA

Palo Alto, CA · On-site +1

$144K - $189K/yr

In this role, you will collaborate with our engineering team to identify performance bottlenecks ... Experience with profiling and debugging tools for GPU applications, such as NVIDIA Nsight.

Senior Platform Engineer

Los Altos, CA · On-site +1

$123K - $169K/yr

... NVIDIA, and Skuchain - and we're looking to grow as we scale Mino to millions of users ... We offer a flexible, remote working environment. You can expect a warm welcome from a friendly and ...

Showing results 21-40

Remote Nvidia Engineering information

What is a remote Nvidia engineer?

A Remote Nvidia Engineer is a professional who works for Nvidia, or with Nvidia technologies, from a location outside of a traditional office setting. These engineers may specialize in areas such as GPU development, AI research, software engineering, or hardware design, and they collaborate with teams virtually. Remote Nvidia Engineers use digital tools to communicate, manage projects, and contribute to cutting-edge technologies in graphics processing, artificial intelligence, and computing platforms. The remote aspect allows for flexible work arrangements and the ability to participate in global projects.

What are some common challenges faced by engineers working remotely for Nvidia, and how can they be overcome?

Remote engineers at Nvidia often encounter challenges related to communication across time zones, staying aligned with fast-paced project developments, and maintaining visibility within distributed teams. To overcome these, it's important to proactively engage in virtual meetings, leverage collaboration tools like Slack and Jira, and regularly update your team on progress. Building strong relationships with peers and seeking out mentorship opportunities can also help remote engineers stay connected and advance within the company.

What are the key skills and qualifications needed to thrive as a remote Nvidia engineer, and why are they important?

To excel as a Remote Nvidia Engineer, you typically need a strong background in computer engineering, programming (e.g., C++, Python), and experience with GPU architectures, often supported by a relevant degree. Familiarity with Nvidia tools like CUDA, cuDNN, and deep learning frameworks, as well as proficiency in remote collaboration platforms, are crucial. Strong problem-solving skills, self-motivation, and effective communication are vital soft skills for working independently and collaborating across distributed teams. These competencies ensure efficient development, troubleshooting, and innovation in Nvidia's complex, high-performance computing environments.

What is the difference between Remote Nvidia Engineering vs Remote Nvidia Data Scientist?

AspectRemote Nvidia EngineeringRemote Nvidia Data Scientist
Required CredentialsBachelor's in Engineering, Computer Science, or related field; experience with GPU programmingBachelor's or higher in Data Science, Statistics, or related; proficiency in machine learning and data analysis
Work EnvironmentDesign, develop, and optimize GPU hardware/software; collaborative teamsAnalyze large datasets, develop models, and generate insights; often cross-functional teams
Employer & Industry UsagePrimarily in hardware, AI, and high-performance computing sectorsPrimarily in AI, analytics, and research sectors

Remote Nvidia Engineering focuses on hardware and software development for GPUs, requiring engineering credentials and technical skills. Remote Nvidia Data Scientists analyze data and build models, requiring expertise in data science. Both roles are remote, but they serve different functions within Nvidia's ecosystem.

What are the most commonly searched types of Nvidia Engineering jobs in California? The most popular types of Nvidia Engineering jobs in California are:
What are popular job titles related to Remote Nvidia Engineering jobs in California? For Remote Nvidia Engineering jobs in California, the most frequently searched job titles are:
What job categories do people searching Remote Nvidia Engineering jobs in California look for? The top searched job categories for Remote Nvidia Engineering jobs in California are:
What cities in California are hiring for Remote Nvidia Engineering jobs? Cities in California with the most Remote Nvidia Engineering job openings:
Infographic showing various Remote Nvidia Engineering job openings in California as of August 2026, with employment types broken down into 88% Full Time, 6% Part Time, 2% Temporary, 3% Contract, and 1% Nights. Highlights an 86% Physical, 5% Hybrid, and 9% Remote job distribution.

Senior Site Reliability Engineer

Andromeda Cluster, Inc

San Francisco, CA • On-site, Remote

$67.25 - $89.25/hr

Full-time

Re-posted 22 days ago


Job description

Senior Site Reliability Engineer
Location: Global Remote / San Francisco • Full-Time
About Andromeda
Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage startups access to the kind of scaled AI infrastructure once reserved only for hyperscalers.
We began with a single managed cluster - but it filled almost instantly. Since then, we've been quietly building the systems, network, and orchestration layer that makes the world's AI infrastructure more accessible.
Today, Andromeda works with leading AI labs, data centers, and cloud providers to deliver compute when and where it's needed most. Our platform routes training and inference jobs across global supply, unlocking flexibility and efficiency in one of the fastest-growing markets on earth.
Our long-term vision is to build the liquidity layer for global AI compute - a marketplace that moves the infrastructure and workloads powering AGI not dissimilar to the flows of capital in the world's financial markets.
We are expanding to new frontiers to find the brightest that work in AI infrastructure, research and engineering.
The Role
This is not a generalist SRE role.
You will design, operate, and debug large-scale GPU infrastructure used for distributed training and inference, working directly with customers pushing the limits of modern AI systems.
We're looking for engineers who have personally run GPU clusters in production, understand the failure modes of distributed training, and can reason about performance from network fabric → kernel → framework.
What You'll Own
  • GPU Cluster Architecture: Design and evolve multi-provider, multi-region GPU compute clusters optimized for large-scale training. Make topology-aware scheduling, networking, and storage decisions that directly impact training throughput and cost efficiency.
  • Customer Technical Partnership: Serve as the primary technical point of contact for customers running large-scale training workloads. Onboard, troubleshoot, and optimize, often in real time.
  • Reliability & Performance Engineering: Define SLOs and error budgets that account for the unique failure modes of GPU infrastructure (ECC errors, NVLink degradation, NCCL timeouts). Own capacity planning across heterogeneous GPU fleets optimized for training throughput.
  • Networking & Fabric Health: Ensure the health and performance of high-speed interconnects (InfiniBand, RoCE, NVLink) that underpin distributed training. Diagnose and resolve fabric-level issues that degrade collective operations.
  • Observability: Build deep visibility into GPU utilization, memory pressure, interconnect throughput, training job performance, and hardware health. Go well beyond standard infrastructure metrics.
  • Automation & Tooling: Build production-grade automation for cluster provisioning, GPU health checks, job scheduling, self-healing, and firmware/driver lifecycle management.
  • Incident Leadership: Lead incident response for complex, multi-layer failures spanning hardware, networking, orchestration, and ML frameworks. Drive blameless postmortems and systemic fixes.

What We're Looking For
  • GPU Systems Expertise: Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience not documentation.
  • High-Performance Networking: Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale.
  • Distributed Training & ML Frameworks: Working knowledge of how large training jobs actually run - NCCL, CUDA, PyTorch distributed, DeepSpeed, Megatron, FSDP, or similar. You don't need to write the models, but you need to understand what's happening at the systems level when a 1,000-GPU training run stalls.
  • Linux & Systems Internals: Expert-level Linux knowledge: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, performance profiling at the syscall and hardware level.
  • Kubernetes & Orchestration: Strong experience running Kubernetes in production with GPU workloads, including device plugins, topology-aware scheduling, multi-cluster federation, and custom operators. Experience with Slurm or other HPC schedulers is equally valued.
  • Automation & Software Engineering: Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts. Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent).
  • Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards.
  • Incident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast.

Strong Candidates May Have
  • Distributed Storage: Experience with high-performance parallel file systems (VAST, Weka, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs.
  • Training Optimization: Experience profiling and optimizing distributed training performance: identifying stragglers, tuning collective communication strategies, improving MFU (Model FLOPs Utilization), and reducing idle GPU time across large runs.
  • Cluster Buildout & Hardware: Experience involved in physical cluster design - rack layout, power/cooling constraints, network topology design, and hardware validation/burn-in at scale.
  • Team Leadership: Experience leading or mentoring a team of infrastructure engineers. We're growing and need people who raise the bar for everyone around them.

Why You'll Love It Here
This is a high-impact, senior builder's role. You'll have significant ownership and autonomy to shape how our systems run at a foundational level, working directly with customers and providers while architecting the infrastructure backbone for reliable, scalable AI compute. You'll influence technical direction and help define what world-class AI infrastructure operations look like.
Andromeda Cluster is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.