1

Server Reliability Engineer Jobs in California (NOW HIRING)

The role We're looking for a world-class Site Reliability Engineer to ensure the reliability ... Experience with bare-metal servers and datacenter operations (PXE/iPXE provisioning, IPMI/BMC, RAID ...

Site Reliability Engineer

San Francisco, CA

$67.25 - $89.25/hr

The role We're looking for a world-class Site Reliability Engineer to ensure the reliability ... Experience with bare-metal servers and datacenter operations (PXE/iPXE provisioning, IPMI/BMC, RAID ...

Site Reliability Engineer

San Francisco, CA · On-site

$67.25 - $89.25/hr

Job Summary : 1872 Consulting is seeking a dynamic Site Reliability Engineer to join their rapidly ... Install physical servers in the data center • * Support other Singular Labs teams as required •

Site Reliability Engineer

Sunnyvale, CA · On-site

$170K - $200K/yr

We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You ... Manage datacenter infrastructure (Linux servers, network devices, databases etc,.) * Improve ...

Site Reliability Engineer

Sunnyvale, CA · On-site

$170K - $200K/yr

We are seeking a talented and motivated Site Reliability Engineer to join our engineering team. You ... Manage datacenter infrastructure (Linux servers, network devices, databases etc,.) * Improve ...

Site Reliability Engineer, GNC

Hawthorne, CA · On-site

$57.75 - $76.75/hr

... and physical servers • Work with SpaceX HPC team to monitor and maintain an HPC cluster ... site reliability or DevOps in lieu of a degree • 1+ years of experience with Linux operating ...

DATABASE RELIABILITY ENGINEER SpaceX is looking for a Database Reliability Engineer with strong ... This individual will need to demonstrate strong knowledge in Windows, server/storage infrastructure ...

DATABASE RELIABILITY ENGINEER SpaceX is looking for a Database Reliability Engineer with strong ... Collaborate with internal stakeholders to optimize schemas, queries, and database/server ...

DATABASE RELIABILITY ENGINEER SpaceX is looking for a Database Reliability Engineer with strong ... Collaborate with internal stakeholders to optimize schemas, queries, and database/server ...

Senior Site Reliability Engineer

Sunnyvale, CA · On-site

$67 - $89/hr

About the Role This Senior Site Reliability Engineer position works on-site out of our Sylmar, CA ... This isn't just about keeping servers up; it's about building and maintaining the resilient ...

Showing results 41-60

Server Reliability Engineer information

What cities in California are hiring for Server Reliability Engineer jobs?

Cities in California with the most Server Reliability Engineer job openings:

Infographic showing various Server Reliability Engineer job openings in California as of August 2026, with employment types broken down into 87% Full Time, 8% Part Time, and 5% Contract. Highlights an 85% Physical, 5% Hybrid, and 10% Remote job distribution.

Senior Site Reliability Engineer

El Segundo, CA • On-site

Varda Space Industries
Guided Missile and Space Vehicle Manufacturing • 51 - 200 employees

$153K - $185K/yr

Full-time

Re-posted 9 days ago


Key responsibilities

  • Deploy, maintain, and operate mission-critical applications and infrastructure supporting spacecraft and company-wide systems.

  • Build and evolve Infrastructure as Code (IaC) frameworks using tools such as Terraform.

  • Build and maintain CI/CD pipelines to enable safe, repeatable, and rapid deployments.


Job description

About This Role 

At Varda Space Industries, we're pushing the boundaries of what's possible in space and materials science - and we're looking for bold engineers to help us get there. As a Senior Site Reliability Engineer, you'll be critical in building, scaling, and maintaining the infrastructure that powers our systems on Earth, in orbit, and everything in between. 

We are looking for an experienced engineer with deep working knowledge of Kubernetes and containerized technologies. You are a hands-on operator and builder who applies first-principles thinking to both software delivery (DevOps) and production reliability (SRE), and thrives in complex, mission-critical environments. 

In this role, you will: 

  • Solve challenging technical problems across a wide range of modern technologies. 
  • Apply a software engineering mindset to automate operations and improve system reliability, scalability, and resilience 
  • Design and build infrastructure that enables rapid development - from cloud-based services to embedded software running on spacecraft. 
  • Shape Varda's infrastructure strategy and drive operational excellence across containerized and modernized environments. 

Responsibilities 

  • Deploy, maintain, and operate mission-critical applications and infrastructure supporting spacecraft and company-wide systems. 
  • Build and evolve Infrastructure as Code (IaC) frameworks using tools such as Terraform 
  • Implement and operate observability systems (metrics, logging, tracing) and actionable alerting. 
  • Build and maintain CI/CD pipelines to enable safe, repeatable, and rapid deployments. 
  • Partner with software and hardware engineers to deliver highly operable, reliable, and scalable systems and pipelines, ensuring they have the tools and infrastructure needed for rapid iteration. 
  • Identify, analyze, and resolve system bottlenecks and reliability risks; perform performance tuning and implement long-term stability improvements. 
  • Respond to and resolve production incidents; perform root cause analysis and drive corrective actions through blameless postmortems. 
  • Rotate through the team's on-call schedule to keep critical systems healthy and responsive. 
  • Must be willing to work extended hours and weekends as needed 
  • Occasionally travel to customer sites and other Varda locations to troubleshoot, deploy, or test critical infrastructure. 

Basic Qualifications 

  • Bachelor's degree in computer science, engineering, or related STEM field with 5+ years of Site Reliability Engineering experience, or 7+ years of progressive experience in DevOps, SRE, or Systems Engineering in lieu of a degree.  
  • Experience with Infrastructure as Code (IaC) using tools like Terraform to automate server provisioning and configuration management 
  • Experience operating Kubernetes or similar container orchestration platforms in production environments. 
  • Experience with Prometheus, Grafana, InfluxDB, or similar technologies. 
  • Knowledge of software-defined networking (VPC, Subnets, Firewalls, VPNs, etc.) 
  • Python, Bash, PowerShell (or similar) scripting experience  
  • Positive and strong communication skills, both written and oral 

Preferred Skills and Experience 

  • Experience in provisioning and managing scalable Azure cloud infrastructure using native tools and best practices 
  • Experience implementing configuration management, provisioning, and workflow automation solutions via Infrastructure as Code, CI/CD, and GitOps (e.g., Ansible, Salt, ArgoCD, etc). 
  • Strong understanding of Linux systems and container runtimes (e.g., containerd, Docker) 
  • Experience with GPU workloads or high-throughput computing. 
  • Hands-on experience operating and optimizing High Performance Computing (HPC) environments, including workload schedulers such as Slurm (e.g., queue/partition design, fair-share scheduling, and cluster resource management). 
  • Experience with hybrid environments (cloud + on-prem or edge systems) 
  • Experience debugging distributed systems at scale (network, storage, latency) 
  • Experience with databases and data modeling 
Pay Range
  • Site Reliability Engineer: $153,000.00 - $185,000.00/per year
  • This role is on-site in El Segundo, CA
  • Leveling and base salary are determined by job-related skills, education level, experience level, and job performance
  • You will be eligible for long-term incentives in the form of stock options and/or long-term cash awards