2

Remote Gpu Programming Jobs in California (NOW HIRING)

Software Engineer

San Diego, CA · On-site +1

$87K - $157K/yr

... ocean remote sensing, and high-performance computing . We're looking for a Software Engineer ... Hands-on experience with GPU/CUDA programming and optimization. * Experience with AI/ML and LLM ...

Remote Commitment: 40 hours/week Role Responsibilities * Guide research and engineering teams to ... Experience writing or optimizing custom GPU kernels using Pallas or Triton . * Demonstrable career ...

next page

Showing results 1-20

Remote Gpu Programming information

What are some common challenges faced by professionals in remote GPU programming roles, and how can they be addressed?

Remote GPU programming roles often involve unique challenges such as managing high-latency connections to remote servers, troubleshooting hardware-specific issues without physical access, and ensuring code compatibility across different GPU architectures. Effective communication with distributed teams is crucial, as is using robust remote debugging tools and version control systems. Staying proactive with documentation and regularly syncing with team members can help address these obstacles and support successful project delivery.

What is remote GPU programming?

Remote GPU programming refers to the practice of developing and running code that utilizes graphics processing units (GPUs) on computers or servers that are accessed over a network, rather than on your local machine. This approach allows developers to leverage powerful, often cloud-based, GPU resources to handle computationally intensive tasks like machine learning, scientific simulations, or rendering without needing specialized hardware themselves. It often involves using remote desktop tools, cloud platforms, or custom APIs to access and manage GPU resources remotely.

What are the key skills and qualifications needed to thrive as a Remote GPU Programmer, and why are they important?

To thrive as a Remote GPU Programmer, you need in-depth knowledge of parallel computing, proficiency in programming languages like C/C++, and experience with GPU architectures, often backed by a degree in computer science or a related field. Familiarity with technical tools such as CUDA, OpenCL, and GPU profiling/debugging systems is commonly required, along with certifications in GPU programming or high-performance computing. Strong problem-solving abilities, self-motivation, and effective remote communication skills help individuals excel in distributed teams. These competencies are crucial for efficiently developing and optimizing GPU-accelerated applications while collaborating across remote environments.
What are the most commonly searched types of Gpu Programming jobs in California? The most popular types of Gpu Programming jobs in California are:
What are popular job titles related to Remote Gpu Programming jobs in California? For Remote Gpu Programming jobs in California, the most frequently searched job titles are:
What job categories do people searching Remote Gpu Programming jobs in California look for? The top searched job categories for Remote Gpu Programming jobs in California are:
What cities in California are hiring for Remote Gpu Programming jobs? Cities in California with the most Remote Gpu Programming job openings:
Infographic showing various Remote Gpu Programming job openings in California as of July 2026, with employment types broken down into 100% Full Time. Highlights an 100% Remote job distribution.
Application Software Engineer, Inference

Application Software Engineer, Inference

SpaceX

Palo Alto, CA • On-site, Remote

$155K - $185K/yr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Re-posted 2 days ago


SpaceX rating

8.8

Company rating: 8.8 out of 10

Based on 146 frontline employees who took The Breakroom Quiz

7th of 61 rated aerospace companies


Job description

APPLICATION SOFTWARE ENGINEER, INFERENCE

The application software team is the central nervous system of SpaceX - we create mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and flight as well as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. We are looking for engineers who treat fellow teammates with fairness, respect, and support.

Our team maintains a high-performance AI inference platform that serves the best models internally at SpaceX to accelerate our most ambitious engineering goals. As part of this effort in Palo Alto, you will design and optimize large-scale model serving systems end-to-end, owning everything from distributed infrastructure to deep low-level optimizations. You will work on systems that deliver reliable, high-throughput inference to power SpaceX's mission-critical applications while maintaining the highest standards of performance and availability.

Aerospace experience is not required to be successful here - rather we look for smart, motivated, respectful, collaborative engineers who love solving problems and want to make an impact on a super inspiring mission. You will have full ownership of challenging problems, working with a team of enthusiastic engineers with diverse perspectives to design and produce solutions that enable SpaceX to achieve its loftiest engineering goals at a rapid pace. The success of the missions at SpaceX depends on the software that you and your team produce.

This role will report through SpaceX Application Software while also working closely with xAI engineering teams. 

RESPONSIBILITIES:

  • Develop highly reliable, high-throughput inference systems that serve the best AI models internally across SpaceX
  • Architect and implement scalable distributed infrastructure for model serving, including load balancing, auto-scaling, batch scheduling, global KV cache, and continuous batching 
  • Optimize latency and throughput of model inference under real production workloads, including low-level GPU kernel work, quantization, speculative decoding, and other acceleration techniques 
  • Build reliable, high-concurrency serving systems with 100% uptime, low tail latency, and excellent observability 
  • Own end-to-end components such as request routing, SDK development, rate limiting, and efficient scaling for internal SpaceX AI inference platforms 
  • Benchmark, fine-tune, and accelerate inference engines (e.g., SGLang, vLLM, TensorRT-LLM) 
  • Develop custom tools for tracing, replaying, and resolving issues across the full stack - from orchestration down to GPU kernels 
  • Create robust CI/CD infrastructure for seamless endpoint deployment, image publishing, and inference engine updates 
  • Collaborate across SpaceXAI teams to integrate inference capabilities into broader systems and workflows 

BASIC QUALIFICATIONS:

  • Bachelor's degree in computer science, engineering, math, or scientific discipline; OR 2+ years of professional experience building software in lieu of a degree
  • Experience in designing, implementing, and maintaining reliable and horizontally scalable distributed systems
  • 1+ years of experience in full stack development or backend development with production systems
  • 1+ years of experience with Rust or C++

PREFERRED SKILLS AND EXPERIENCE:

  • Experience with LLM inference engines and serving frameworks (e.g., SGLang, vLLM, Triton, TensorRT-LLM) 
  • Deep low-level systems programming and optimizations: GPU kernels, code generation, batching, caching, parallelism, quantization, and speculative decoding 
  • Experience with large-scale, high-concurrency production serving systems 
  • Knowledge of service observability and reliability best practices 
  • Experience operating commonly used databases such as PostgreSQL, ClickHouse, or MongoDB 
  • Experience designing or building with agent SDKs and agent orchestration frameworks 
  • Experience with Docker, Kubernetes, and containerized applications 
  • Expert knowledge of gRPC (unary, response streaming, bi-directional streaming, REST mapping) 
  • Programming experience in Python, Go, or similar languages 
  • Experience with version control, continuous integration, continuous delivery, build systems, and monitoring 
  • Expertise in profiling and improving application performance 

ADDITIONAL REQUIREMENTS:

  • You may be asked to work extended hours/weekends dependent on launch cadence and platform demands 
  • This role requires you to be onsite in Palo Alto. Remote and/or hybrid work will not be considered 

COMPENSATION AND BENEFITS:
Pay Range:
Software Engineer/Level I: $135,000.00 - $160,000.00/per year
Software Engineer/Level II: $155,000.00 - $185,000.00/per year

Your actual level and base salary will be determined on a case-by-case basis and may vary based on the following considerations: job-related knowledge and skills, education, and experience.

Base salary is just one part of your total rewards package at SpaceX. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan. You will also receive access to comprehensive medical, vision, and dental coverage, access to a 401(k) retirement plan, short and long-term disability insurance, life insurance, paid parental leave, and various other discounts and perks. You may also accrue 3 weeks of paid vacation and will be eligible for 10 or more paid holidays per year. Employees accrue paid sick leave pursuant to Company policy which satisfies or exceeds the accrual, carryover, and use requirements of the law.


What SpaceX employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom