2

Remote Infrastructure Engineer Jobs in California

Staff AI Infrastructure Engineer

Redwood City, CA ยท On-site +1

$131K - $172K/yr

Achieving that requires training frontier-scale AI biology models, and that demands reliable, high-performance compute infrastructure. This is production engineering work at a frontier AI lab, with ...

Showing results 21-40

Remote Infrastructure Engineer information

What is a remote infrastructure engineer?

A Remote Infrastructure Engineer is responsible for designing, managing, and maintaining an organization's IT infrastructure, including servers, networks, cloud environments, and security systems, while working remotely. They ensure system reliability, optimize performance, troubleshoot issues, and implement new technologies to support business operations. This role requires expertise in networking, cloud computing, automation, and cybersecurity, often working with tools like AWS, Azure, VMware, and monitoring systems. Effective communication and problem-solving skills are essential, as they collaborate with teams to resolve technical challenges.

What does a remote infrastructure engineer do?

As a Remote Infrastructure Engineer, your daily responsibilities generally include monitoring system performance, deploying updates and patches, managing cloud resources, and troubleshooting network or server issues. You'll often work closely with development and security teams to ensure optimal infrastructure performance, automate routine processes, and uphold security standards. Collaboration through project management platforms and remote communication tools is common to keep the team aligned. The role frequently involves both scheduled maintenance and responding to urgent incidents, offering a dynamic and impactful work environment.

What skills and qualifications are needed to be a remote infrastructure engineer?

To thrive as a Remote Infrastructure Engineer, you need a strong background in network administration, cloud computing, and server management, often paired with a degree in computer science or related field. Familiarity with tools like AWS, Azure, Docker, Kubernetes, and certifications such as AWS Certified Solutions Architect or CompTIA Network+ are typically expected. Excellent problem-solving abilities, proactive communication, and strong organizational skills help you manage tasks independently and collaborate across dispersed teams. These capabilities are crucial for maintaining reliable, secure infrastructure and supporting seamless remote operations for organizations.

What are the most commonly searched types of Infrastructure Engineer jobs in California?

The most popular types of Infrastructure Engineer jobs in California are:

What job categories do people searching Remote Infrastructure Engineer jobs in California look for?

The top searched job categories for Remote Infrastructure Engineer jobs in California are:

What cities in California are hiring for Remote Infrastructure Engineer jobs?

Cities in California with the most Remote Infrastructure Engineer job openings:

Infographic showing various Remote Infrastructure Engineer job openings in California as of August 2026, with employment types broken down into 95% Full Time, 2% Part Time, and 3% Contract. Highlights an 85% Physical, 6% Hybrid, and 9% Remote job distribution.

Reinforcement Learning Infrastructure Engineer

Elorian

Palo Alto, CA โ€ข On-site, Remote

$200K - $400K/yr

Full-time

Medical, Dental, Vision, PTO

Posted 29 days ago


Job description

About Us
We are a well-funded, early-stage AI lab focused on building the next generation of frontier multimodal AI models. Founded by former DeepMind researchers, including Andrew Dai, who was previously a leader on Gemini. Our team currently consists of 20 world-class scientists and engineers. We recently raised $55M in seed funding from Striker Ventures, Menlo Ventures, Altimeter Capital, and NVIDIA. We are tackling some of the hardest problems in artificial intelligence, and we are growing fast.
The Role
We're looking for an infrastructure engineer to design and build the core systems behind how we train our models with reinforcement learning (RL).
You'll own the training infrastructure end to end, from rollout and reward pipelines to orchestration, reliability, and observability. The work spans both the algorithmic side of RL and the systems reality of running distributed training at scale, and you'll partner closely with our research team to keep RL training fast, stable, and dependable for the multimodal, visual reasoning models at the center of our work.
What You Will Do
  • Design, build, and optimize the infrastructure that powers our large-scale RL and post-training workloads
  • Improve the reliability, scalability, and throughput of distributed RL training pipelines
  • Build actor-learner architectures and orchestrate environment rollouts at scale
  • Develop monitoring and observability tools that ensure high uptime, debuggability, and reproducibility across RL systems
  • Collaborate with researchers to translate algorithmic ideas into production-grade training pipelines
  • Improve GPU utilization and training throughput across the cluster

What We're Looking For
Minimum qualifications:
  • 3+ years of distributed systems experience, including building or optimizing large-scale RL training pipelines (PPO, GRPO, or similar on-policy methods)
  • Experience with actor-learner architectures and environment rollout orchestration at scale
  • Strong Python skills, plus PyTorch or JAX
  • Experience with async training infrastructure, replay buffers, or simulation-based environment frameworks
  • Multi-node GPU orchestration experience (Ray, SLURM, or Kubernetes)
  • A track record of improving training throughput and GPU utilization at scale
  • Strong engineering skills; ability to contribute performant, maintainable code and debug in complex codebases

Preferred qualifications (strong candidates may have some, not all):
  • Experience with multimodal or agentic RL environments
  • Experience with RLHF or reward modeling pipelines
  • A self-directed builder who moves quickly and works across teams in an early-stage setting

Logistics
Location: This role is based on-site in Palo Alto, California.
Compensation: Depending on background, skills, and experience, the expected annual base salary range for this position is $200,000 - $400,000 USD, plus equity and benefits.
Visa sponsorship: We sponsor work visas. We can't promise every case will succeed, but for the right person we'll work through the process with you.
Benefits: We offer health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.
Elorian AI is an equal opportunity employer. We are committed to building a diverse team and inclusive environment.