2

Overnight Site Reliability Engineer Remote Jobs in Texas

Cloud Site Reliability Engineer

Dallas, TX ยท Remote

$56.50 - $75/hr

Stefanini is looking for Cloud Site Reliability Engineer - Remote For quick apply, please contact Sudhanshu.Shrivastava; Ph: 248 582 6510 Sudhanshu.shrivastava@stefanini.com W2 Candidates only! About ...

Cloud Site Reliability Engineer

Dallas, TX ยท Remote

$56.50 - $75/hr

Stefanini is looking for Cloud Site Reliability Engineer - Remote For quick apply, please contact Sudhanshu.Shrivastava; Ph: 248 582 6510 Sudhanshu.shrivastava@stefanini.com W2 Candidates only! About ...

Cloud Site Reliability Engineer

Dallas, TX ยท Remote

$56.50 - $75/hr

Stefanini is looking for Cloud Site Reliability Engineer - Remote For quick apply, please contact Sudhanshu.Shrivastava; Ph: 248 582 6510 Sudhanshu.shrivastava@stefanini.com W2 Candidates only! About ...

Site Reliability Engineer

Westlake, TX ยท On-site +1

$54.75 - $72.75/hr

We are currently seeking a Site Reliability Engineer to join our team in Westlake, Texas (US-TX), ... The starting pay range for this remote role is 85,000-140,000 This range reflects the minimum and ...

Site Reliability Engineer

Austin, TX ยท On-site +1

$56.50 - $75/hr

S. Citizenship / No clearance needed / 100% remote within the US Staff Site Reliability Engineer / Cloud SME Location: 100% remote in the continental US Type: Long-term contract (3+ years) Role ...

Senior Site Reliability Engineer II

Allen, TX ยท On-site +1

$125K - $209K/yr

If not, this role is fully remote. We do not restrict applicants based on job site or posting location. Job Title: Senior Site Reliability Engineer (SRE) Location: Open (U.S.-based). No job site ...

Senior Site Reliability Engineer II

Dallas, TX ยท On-site +1

$125K - $209K/yr

If not, this role is fully remote. We do not restrict applicants based on job site or posting location. Job Title: Senior Site Reliability Engineer (SRE) Location: Open (U.S.-based). No job site ...

Sr. SRE Platform Software Engineer

Austin, TX ยท On-site +1

$56.50 - $75/hr

To learn more, visit Position Overview Build and operate one or more bounded contexts of the NeoCloud SRE platform - the multi-region substrate that observes, protects, and operates a GPU rental ...

Site Reliability Engineer II

Austin, TX ยท On-site +1

$98K - $138K/yr

The Site Reliability Engineer II will be responsible for supporting, enhancing, and maintaining ... Wellness initiatives #BI-Remote Internal Employees - R365 is committed to growing talent from ...

Senior Site Reliability Engineer

Austin, TX ยท On-site +1

$56.50 - $75/hr

The role We're looking for a Senior SRE to own the reliability, scalability, and operational posture of Satsuma's multi-cloud infrastructure. You'll be the person who keeps things running, builds the ...

Site Reliability Engineer

Windcrest, TX ยท Remote

$51.50 - $68.25/hr

US REMOTE** **WORK FROM HOME** Rackspace Professional Services is helping customers solve business challenges using Kubernetes. RPS is seeking Kubernetes DevOps engineers who can help guide clients ...

next page

Showing results 1-20

Overnight Site Reliability Engineer Remote information

What is an overnight site reliability engineer?

An Overnight Site Reliability Engineer (SRE) is a professional responsible for ensuring the reliability, performance, and uptime of software systems during overnight or off-peak hours, typically working remotely. Their main tasks include monitoring system health, responding to incidents, troubleshooting outages, and implementing fixes to maintain service availability. SREs also work to automate processes, improve system resilience, and collaborate with other engineering teams to prevent future issues. Working overnight ensures that critical systems remain operational and issues are addressed promptly, even outside of standard business hours.

What skills and qualifications are needed to be an overnight site reliability engineer?

To thrive as an Overnight Site Reliability Engineer (Remote), you need strong expertise in systems administration, incident response, automation, and a solid background in computer science or related fields. Proficiency with monitoring tools (like Prometheus or Datadog), cloud platforms (such as AWS or GCP), scripting languages (Python, Bash), and certifications like AWS Certified SysOps Administrator are highly beneficial. Exceptional problem-solving skills, attention to detail, and effective remote communication help you excel in high-pressure overnight scenarios. These skills ensure system reliability, minimize downtime, and maintain seamless operations during critical off-hours.

What are the unique challenges and expectations for an overnight site reliability engineer working remotely?

As an Overnight Site Reliability Engineer working remotely, you'll often handle critical incidents that arise outside of standard business hours, so strong problem-solving skills and the ability to work independently are crucial. Communication is key, as you'll need to coordinate with team members in different time zones and document incidents clearly for seamless handoffs. You may also be tasked with proactive monitoring and maintenance activities during quieter periods, making self-motivation and attention to detail especially important. The role offers valuable exposure to high-impact issues and can accelerate your expertise in incident management and system reliability.

What is the difference between Overnight Site Reliability Engineer Remote vs Overnight DevOps Engineer Remote?

AspectOvernight Site Reliability Engineer RemoteOvernight DevOps Engineer Remote
Primary FocusEnsuring system reliability, uptime, and incident responseAutomating deployment, integration, and infrastructure management
Required SkillsMonitoring, incident management, scripting, system troubleshootingCI/CD pipelines, automation, cloud platforms, scripting
Work EnvironmentRemote, on-call shifts, collaboration with SRE teamsRemote, development and deployment focus, collaboration with development teams
CertificationsLinux, AWS, Google Cloud, or Azure certifications often preferredCloud certifications, Docker, Kubernetes, CI/CD tools

While both roles involve working remotely and require cloud and scripting skills, the Overnight Site Reliability Engineer Remote primarily focuses on maintaining system reliability and incident response, whereas the Overnight DevOps Engineer Remote emphasizes automation, deployment, and infrastructure management. Understanding these differences helps candidates align their skills with the right role.

What are the most commonly searched types of Site Reliability Engineer Remote jobs in Texas?

The most popular types of Site Reliability Engineer Remote jobs in Texas are:

What job categories do people searching Overnight Site Reliability Engineer Remote jobs in Texas look for?

The top searched job categories for Overnight Site Reliability Engineer Remote jobs in Texas are:

What cities in Texas are hiring for Overnight Site Reliability Engineer Remote jobs?

Cities in Texas with the most Overnight Site Reliability Engineer Remote job openings:

Cloud Site Reliability Engineer

Stefanini Group

Dallas, TX โ€ข Remote

$56.50 - $75/hr

Contractor

Re-posted 18 days ago


Job description

Stefanini Group is hiring!
Stefanini is looking forย Cloud Site Reliability Engineer - Remote
For quick apply, please contact Sudhanshu.Shrivastava; Ph: 248 582 6510
Sudhanshu.shrivastava@stefanini.com
ย 
W2 Candidates only!
ย 
About the Opportunity:
As a Senior Cloud Engineer in the Cloud SRE team, you will be responsible for designing and developing cloud solutions and engineering reliability tools for the Cloud Foundation Services (CFS) platform in the Infrastructure, Platforms & Operations organization. You will apply software engineering practices to build scalable, reusable solutions and utilities that enhance platform reliability.
ย 
Responsibilities:
What Will Be Expected of You:
  • Design, develop, and maintain reliability solutions and SRE utilities to reduce toil, improve cloud platform reliability, and industrialize SRE practices across the system
  • Build and optimize Infrastructure as Code (IaC) using Terraform to manage AWS resources related to SRE solutions, incorporating cost-efficient design principles
  • Develop CI/CD pipelines and automated testing to ensure code quality, reliability, and rapid delivery of the solutions
  • Define SRE standards, best practices, and guidelines for adoption across teams; establish SRE metrics like SLI, SLOs, etc.
  • Apply software engineering best practices including version control, code reviews, test-driven development, and documentation to all development
  • Participate in incident management and on-call rotation, providing technical support for SRE tools, troubleshooting production issues, and collaborating with teams to reduce incident recurrence through proactive detection and pattern analysis
  • Stay current with emerging AWS services, SRE methodologies, and cloud-native development technologies, and drive adoption of innovative solutions
  • Collaborate within Agile and Scaled Agile frameworks with cross-functional teams to deliver integrated cloud automation solutions
  • Produce clear, blameless postmortems with actionable items and documented failure scenarios
ย 
ย #LI-SS3
#LI-REMOTE

Qualifications:
  • Bachelor's degree in computer science, Information Systems, or equivalent background or equivalent experience
  • 7+ years of extensive experience in software development with focus on reliability and platform engineering
  • 5+ Years of advanced Python development skills with proven experience building enterprise-grade, highly available tools, APIs, and utilities
  • 3+ years of hands-on experience developing solutions in AWS environments with deep understanding of core services (EC2, VPC, S3, Lambda, IAM, CloudFormation, EventBridge, Step Functions etc.) and resource cost optimization
  • 3+ years of experience applying SRE principles including observability, toil automation, SLIs/SLOs and reliability engineering
  • Expert-level proficiency with Infrastructure as Code (IaC) using Terraform, including module development and state management
  • Strong experience with CI/CD pipelines, automated testing frameworks, and DevOps practices
  • Experience with observability tools and practices including Grafana, AWS CloudWatch, AWS Canary
  • Experience defining, implementing, and managing SLOs/SLIs and error budgets; familiarity with conducting RCAs and producing postmortem documentation
  • Working experience in Agile and Scaled Agile environments and familiarity with ITSM processes (incident, change, and problem management), resilience testing and chaos engineering practices
  • Experience with GoLang or additional programming languages is a plus
ย 
ย 
Stefanini takes pride in hiring top talent and developing relationships with our future employees. Our talent acquisition teams will never make an offer of employment without having a phone conversation with you. Those face-to-face conversations will involve a description of the job for which you have applied. We also speak with you about the process including interviews and job offers.
ย 
About Stefanini Group:
The Stefanini Group is a global provider of offshore, onshore and near shore outsourcing, IT digital consulting, systems integration, application, and strategic staffing services to Fortune 1000 enterprises around the world. Our presence is in countries like the Americas, Europe, Africa, and Asia, and more than four hundred clients across a broad spectrum of markets, including financial services, manufacturing, telecommunications, chemical services, technology, public sector, and utilities. Stefanini is a CMM level 5, IT consulting company with a global presence. We are CMM Level 5 company
Education:Bachelor (BA, BS...)Employment Type: CONTRACTOR