2

Overnight Site Reliability Engineer Remote Jobs in California

Senior Site Reliability Engineer

San Diego, CA · Remote

$60.50 - $80.50/hr

This is a remote, contract opportunity for a project Arctiq is delivering for a client. Candidates ... The Senior Site Reliability Engineer is a technical leader responsible for architecting the ...

Site Reliability Engineer

Palo Alto, CA · On-site +1

$165K - $190K/yr

About the DevOps / SRE Team The DevOps/SRE team at Obsidian ensures that engineering excellence translates into stable, scalable, and high-performing production systems. We work closely with ...

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... Wellness initiatives #BI-Remote Internal Employees - R365 is committed to growing talent from ...

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... Wellness initiatives #BI-Remote Internal Employees - R365 is committed to growing talent from ...

Senior Site Reliability Engineer

Glendale, CA · On-site +1

$60.50 - $80.25/hr

Our Site Reliability and Infrastructure Engineering team centralizes the concerns of measurement and guidance so every engineer can improve availability and efficiency in their own area of the ...

next page

Showing results 1-20

Overnight Site Reliability Engineer Remote information

What is an overnight site reliability engineer?

An Overnight Site Reliability Engineer (SRE) is a professional responsible for ensuring the reliability, performance, and uptime of software systems during overnight or off-peak hours, typically working remotely. Their main tasks include monitoring system health, responding to incidents, troubleshooting outages, and implementing fixes to maintain service availability. SREs also work to automate processes, improve system resilience, and collaborate with other engineering teams to prevent future issues. Working overnight ensures that critical systems remain operational and issues are addressed promptly, even outside of standard business hours.

What are the unique challenges and expectations for an overnight site reliability engineer working remotely?

As an Overnight Site Reliability Engineer working remotely, you'll often handle critical incidents that arise outside of standard business hours, so strong problem-solving skills and the ability to work independently are crucial. Communication is key, as you'll need to coordinate with team members in different time zones and document incidents clearly for seamless handoffs. You may also be tasked with proactive monitoring and maintenance activities during quieter periods, making self-motivation and attention to detail especially important. The role offers valuable exposure to high-impact issues and can accelerate your expertise in incident management and system reliability.

What skills and qualifications are needed to be an overnight site reliability engineer?

To thrive as an Overnight Site Reliability Engineer (Remote), you need strong expertise in systems administration, incident response, automation, and a solid background in computer science or related fields. Proficiency with monitoring tools (like Prometheus or Datadog), cloud platforms (such as AWS or GCP), scripting languages (Python, Bash), and certifications like AWS Certified SysOps Administrator are highly beneficial. Exceptional problem-solving skills, attention to detail, and effective remote communication help you excel in high-pressure overnight scenarios. These skills ensure system reliability, minimize downtime, and maintain seamless operations during critical off-hours.

What is the difference between Overnight Site Reliability Engineer Remote vs Overnight DevOps Engineer Remote?

AspectOvernight Site Reliability Engineer RemoteOvernight DevOps Engineer Remote
Primary FocusEnsuring system reliability, uptime, and incident responseAutomating deployment, integration, and infrastructure management
Required SkillsMonitoring, incident management, scripting, system troubleshootingCI/CD pipelines, automation, cloud platforms, scripting
Work EnvironmentRemote, on-call shifts, collaboration with SRE teamsRemote, development and deployment focus, collaboration with development teams
CertificationsLinux, AWS, Google Cloud, or Azure certifications often preferredCloud certifications, Docker, Kubernetes, CI/CD tools

While both roles involve working remotely and require cloud and scripting skills, the Overnight Site Reliability Engineer Remote primarily focuses on maintaining system reliability and incident response, whereas the Overnight DevOps Engineer Remote emphasizes automation, deployment, and infrastructure management. Understanding these differences helps candidates align their skills with the right role.

What are the most commonly searched types of Site Reliability Engineer Remote jobs in California? The most popular types of Site Reliability Engineer Remote jobs in California are:
What are popular job titles related to Overnight Site Reliability Engineer Remote jobs in California? For Overnight Site Reliability Engineer Remote jobs in California, the most frequently searched job titles are:
What job categories do people searching Overnight Site Reliability Engineer Remote jobs in California look for? The top searched job categories for Overnight Site Reliability Engineer Remote jobs in California are:
What cities in California are hiring for Overnight Site Reliability Engineer Remote jobs? Cities in California with the most Overnight Site Reliability Engineer Remote job openings:

Senior Site Reliability Engineer

Arctiq, Inc.

San Diego, CA • Remote

$60.50 - $80.50/hr

Contractor

Re-posted 25 days ago


Job description

Company Overview:
Arctiq is a global, intelligence-driven technology services company delivering professional and managed services across Hybrid Cloud Infrastructure, Networking & Connected Experiences, Cybersecurity, Data & AI, Autonomous Operations & Intelligence, and Enterprise Service Management. We help organizations operate, secure, and modernize complex environments by unifying infrastructure, networking, data, security, automation, and observability under a single, integrated operating model. Our work focuses on helping customers reduce operational friction, improve resilience, and make better, faster decisions as their environments evolve. Arctiq builds on decades of industry expertise and a customer-centric ethos to deliver exceptional value to clients across diverse industries.


This is a remote, contract opportunity for a project Arctiq is delivering for a client. Candidates must have or be able to obtain a Secret Clearance.


This position requires U.S. citizenship due to government security clearance and/or federal contract requirements


This is a remote opportunity with preference to candidates located in San Diego, CA, Norfolk, VA or Charleston, SC


Position Overview:

The Senior Site Reliability Engineer is a technical leader responsible for architecting the reliability strategy for large-scale, distributed government systems. You will lead the implementation of the SRE framework, driving the adoption of SLO-based management and advanced automation. As a subject matter expert, you will mentor mid-level engineers and interface with government stakeholders to ensure system resilience and performance meet mission requirements.



Responsibilities:


  • Reliability Architecture: Define the strategy for Service Level Objectives (SLOs) and Error Budgets. Design complex telemetry pipelines for full-stack observability.
  • Strategic Automation: Design and govern the enterprise Infrastructure as Code (IaC) standards. Develop custom tooling to automate complex recovery procedures and system scaling.
  • Incident Command: Act as the Incident Commander for major system outages, leading the technical response and directing the Root Cause Analysis (RCA) process.
  • Security & Compliance: Lead the integration of security-as-code within DevSecOps pipelines, ensuring full compliance with RMF and NIST 800-53 standards.
  • Mentorship: Provide technical guidance and mentorship to Mid-Level SREs and developers, fostering a culture of reliability across the organization.


Qualifications:


  • 7+ years of experience in SRE or DevOps, with significant experience in distributed systems.
  • Expertise in Go, Python, or Java and advanced knowledge of Linux internals.
  • Extensive experience managing production Kubernetes environments and complex cloud architectures.
  • Proven track record of defining and meeting SLOs for high-availability systems.
  • Experience navigating government Risk Management Framework (RMF) processes.
  • Education: Bachelor's or Master's degree in Computer Science or Engineering.
  • Certifications: CKA (Certified Kubernetes Administrator) and industry observability certification preferred
  • Strong Communication
  • Comfortable Interacting with Leadership
  • Promoting Best Practices
  • Understanding Operations and Procedures to help lead the team
  • 7+ Years of Leadership Experience
  • Executive Presence


Arctiq is an equal opportunity employer. If you need any accommodations or adjustments throughout the interview process and beyond, please let us know. We celebrate our inclusive work environment and welcome members of all backgrounds and perspectives to apply.

We thank you foryour interest in joining the Arctiq team! While we welcome all applicants, only those who are selected for an interview will be contacted.