2

Overnight Site Reliability Engineer Remote Jobs in California

Staff Site Reliability Engineer

Santa Barbara, CA · On-site +1

$63.50 - $84.25/hr

PayJunction is seeking a Staff Site Reliability Engineer to own the reliability, security, and ... Opportunity for remote, in-office, or hybrid work Office Environment * The opportunity to choose ...

Staff Site Reliability Engineer

Mountain View, CA · On-site +1

$67.25 - $89.25/hr

Certain roles -- such as field-based sales or other remote-by-design positions -- may have ... Role Overview The Staff Site Reliability Engineer, Infrastructure role is building a high-scale ...

Staff Site Reliability Engineer

San Mateo, CA · On-site +1

$200K - $285K/yr

We are looking for a hands-on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This role is focused on owning production infrastructure ...

AI-First SRE/DevOps Engineer

San Jose, CA · On-site +1

$66.75 - $88.75/hr

US (Remote or HQ Hybrid) Job Type: Full-time Axiad is seeking a skilled AI-First SRE/DevOps Engineer with 5-8 years of hands-on infrastructure and platform engineering experience to help build and ...

Showing results 21-40

Overnight Site Reliability Engineer Remote information

What is an overnight site reliability engineer?

An Overnight Site Reliability Engineer (SRE) is a professional responsible for ensuring the reliability, performance, and uptime of software systems during overnight or off-peak hours, typically working remotely. Their main tasks include monitoring system health, responding to incidents, troubleshooting outages, and implementing fixes to maintain service availability. SREs also work to automate processes, improve system resilience, and collaborate with other engineering teams to prevent future issues. Working overnight ensures that critical systems remain operational and issues are addressed promptly, even outside of standard business hours.

What skills and qualifications are needed to be an overnight site reliability engineer?

To thrive as an Overnight Site Reliability Engineer (Remote), you need strong expertise in systems administration, incident response, automation, and a solid background in computer science or related fields. Proficiency with monitoring tools (like Prometheus or Datadog), cloud platforms (such as AWS or GCP), scripting languages (Python, Bash), and certifications like AWS Certified SysOps Administrator are highly beneficial. Exceptional problem-solving skills, attention to detail, and effective remote communication help you excel in high-pressure overnight scenarios. These skills ensure system reliability, minimize downtime, and maintain seamless operations during critical off-hours.

What are the unique challenges and expectations for an overnight site reliability engineer working remotely?

As an Overnight Site Reliability Engineer working remotely, you'll often handle critical incidents that arise outside of standard business hours, so strong problem-solving skills and the ability to work independently are crucial. Communication is key, as you'll need to coordinate with team members in different time zones and document incidents clearly for seamless handoffs. You may also be tasked with proactive monitoring and maintenance activities during quieter periods, making self-motivation and attention to detail especially important. The role offers valuable exposure to high-impact issues and can accelerate your expertise in incident management and system reliability.

What is the difference between Overnight Site Reliability Engineer Remote vs Overnight DevOps Engineer Remote?

AspectOvernight Site Reliability Engineer RemoteOvernight DevOps Engineer Remote
Primary FocusEnsuring system reliability, uptime, and incident responseAutomating deployment, integration, and infrastructure management
Required SkillsMonitoring, incident management, scripting, system troubleshootingCI/CD pipelines, automation, cloud platforms, scripting
Work EnvironmentRemote, on-call shifts, collaboration with SRE teamsRemote, development and deployment focus, collaboration with development teams
CertificationsLinux, AWS, Google Cloud, or Azure certifications often preferredCloud certifications, Docker, Kubernetes, CI/CD tools

While both roles involve working remotely and require cloud and scripting skills, the Overnight Site Reliability Engineer Remote primarily focuses on maintaining system reliability and incident response, whereas the Overnight DevOps Engineer Remote emphasizes automation, deployment, and infrastructure management. Understanding these differences helps candidates align their skills with the right role.

What are the most commonly searched types of Site Reliability Engineer Remote jobs in California?

The most popular types of Site Reliability Engineer Remote jobs in California are:

What are popular job titles related to Overnight Site Reliability Engineer Remote jobs in California?

For Overnight Site Reliability Engineer Remote jobs in California, the most frequently searched job titles are:

What job categories do people searching Overnight Site Reliability Engineer Remote jobs in California look for?

The top searched job categories for Overnight Site Reliability Engineer Remote jobs in California are:

What cities in California are hiring for Overnight Site Reliability Engineer Remote jobs?

Cities in California with the most Overnight Site Reliability Engineer Remote job openings:

Site Reliability Engineer II, tvScientific

Pinterest

San Francisco, CA • On-site, Remote

$67.25 - $89.25/hr

Full-time

Re-posted 15 days ago


Job description

About tvScientific

tvScientific is the first and only CTV advertising platform purpose-built for performance marketers. We leverage massive data and cutting-edge science to automate and optimize TV advertising to drive business outcomes. Our solution combines media buying, optimization, measurement, and attribution in one, efficient platform. Our platform is built by industry leaders with a long history in programmatic advertising, digital media, and ad verification who have now purpose-built a CTV performance platform advertisers can trust to grow their business.

We are seeking a Site Reliability Engineer to help operate, scale, and continuously improve a cloud-native platform built on AWS, Kubernetes/EKS, and ArgoCD-driven GitOps workflows. This role will contribute to improving the reliability, scalability, automation, observability, and operational maturity of our infrastructure and delivery ecosystem. The ideal candidate is a hands-on engineer with solid production experience and a strong foundation in building and supporting resilient platforms using infrastructure as code, automation, and modern Kubernetes operational practices.

What you'll do:

  • Ensuring the reliability, availability, and performance of production infrastructure and platform services
  • Operating and scaling Kubernetes platforms, including governance and support for multi-tenant workloads
  • Managing GitOps-based deployment workflows using ArgoCD and Helm
  • Supporting infrastructure provisioning and change management through Terraform/Terragrunt
  • Building and supporting CI/CD automation and deployment workflows using GitHub Actions
  • Participating in incident response, root cause analysis, and post-incident improvement initiatives
  • Reducing operational toil through scripting, tooling, and process automation
  • Advancing observability practices across logs, metrics, traces, dashboards, and alerting
  • Supporting secure secrets integration, IAM-aware operations, and platform guardrails
  • Partnering closely with application, security, and platform teams to improve reliability and delivery outcomes

What we're looking for:

  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure
  • Strong hands-on experience operating AWS in production environments
  • Good expertise in Kubernetes, including cluster operations, troubleshooting, workload reliability, and platform administration
  • Experience with Kubernetes multi-tenancy, including namespaces, RBAC, quotas, policies, and tenant isolation patterns
  • Experience implementing and operating ArgoCD within a GitOps delivery model
  • Strong hands-on experience with Helm
  • Experience with Terraform/Terragrunt for infrastructure provisioning and environment management
  • Solid scripting and automation skills using Bash and/or Python
  • Experience building, maintaining, or supporting CI/CD pipelines, ideally using GitHub Actions
  • Strong troubleshooting skills across Linux, containers, IAM, networking, and distributed systems
  • Experience with monitoring, alerting, and observability in production environments
  • Demonstrated ownership mindset with experience handling incidents and resolving production issues
  • Strong collaboration and communication skills, with the ability to work effectively across engineering, security, and platform teams
  • Bachelor's degree in computer science, engineering, a related field or equivalent experience
  • Demonstrated ability to use AI to improve speed and quality in your day-to-day workflow for relevant outputs
  • Strong track record of critical evaluation and verification of AI-assisted work (e.g., testing, source-checking, data validation, peer review)
  • High integrity and ownership: you protect sensitive data, avoid over-reliance on AI, and remain accountable for final decisions and deliverables

In-Office Requirement Statement:

  • We recognize that the ideal environment for work is situational and may differ across departments. What this looks like day-to-day can vary based on the needs of each organization or role.


Relocation Statement:

  • This position is not eligible for relocation assistance. Visit our PinFlex page to learn more about our working model.

#LI-SM4

#LI-REMOTE