1

Services Reliability Engineer Jobs (NOW HIRING)

They are seeking an experienced Senior Site Reliability Engineer to enhance the reliability ... Services contract. Responsibilities : • Monitor and maintain system reliability, availability ...

The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held ... This SRE will configure, tune, and fix multi-tiered systems to achieve optimal application ...

Own service reliability and operational health-establish and maintain SLOs/SLIs, design monitoring ... Partner with development and engineering teams to evaluate deployment readiness, support deployment ...

Seattle, Washington, United States Software and Services The Apple Services Engineering team (ASE ... This SRE will configure, tune, and fix multi‑tiered systems to achieve optimal application ...

SRE

Atlanta, GA · On-site

$54.75 - $72.75/hr

Architect and design highly available, scalable, secure, and cost-effective infrastructure and application patterns on AWS * and evangelize SRE best practices, standards, and blueprints for service ...

SRE with FedRAMP

$58.25 - $77.50/hr

Rootshell Enterprise Technologies Inc. is a recognized provider of professional IT Consulting services in the US. We are actively seeking SRE with FedRAMP for one of our client, Please share your ...

Site Reliability Engineer (SRE)

Charlotte, NC · On-site

$55.75 - $74/hr

Generate and maintain Critical User Journeys (CUJs), Service Level Indicators (SLIs), and Service Level Objectives (SLOs). * Conduct blameless post-mortems and manage error budgets. * Develop and ...

Waymo's software reliability engineers (SRE) are responsible for the stable operation of Waymo ... Collaborate with other engineers to build reliable systems for the ride hailing service, real time ...

Respond to and resolve system outages, impairments, and service disruptions while coordinating with ... Expert knowledge of site reliability engineering practices, system monitoring, incident management ...

Respond to and resolve system outages, impairments, and service disruptions while coordinating with ... Expert knowledge of site reliability engineering practices, system monitoring, incident management ...

Reliability Engineer III

Clifton, NY · On-site

$80K - $100K/yr

Position Summary The Field Service Reliability Engineer is a key member of NECI's Reliability Solutions team, responsible for delivering Reliability Engineering Services to industrial clients ...

Reliability Engineer III

Clifton Park, NY · On-site

$99K - $124K/yr

Position Summary The Field Service Reliability Engineer is a key member of NECIs Reliability Solutions team, responsible for delivering Reliability Engineering Services to industrial clients ...

Site Reliability Engineer III Location: Pennington, NJ Duration: Contract - 7 months Pay Range: $73 ... Founded in 1998, BCforward has grown with our customers needs into a full-service business ...

New

Showing results 21-40

Services Reliability Engineer information

See salary details

$61K

$118K

$141K

How much do services reliability engineer jobs pay per year?

As of Sep 14, 2026, the average yearly pay for services reliability engineer in the United States is $117,973.00, according to ZipRecruiter salary data. Most workers in this role earn between $102,500.00 and $129,000.00 per year, depending on experience, location, and employer.

What cities are hiring for Services Reliability Engineer jobs?

Cities with the most Services Reliability Engineer job openings:

What states have the most Services Reliability Engineer jobs?

States with the most job openings for Services Reliability Engineer jobs include:

What are popular job titles related to Services Reliability Engineer jobs?

For Services Reliability Engineer jobs, the most frequently searched job titles are:

Senior Reliability Engineer

Washington, DC • On-site

Barbaricum
National Security • 201 - 500 employees

Full-time

Re-posted 28 days ago


Job description

Job Summary:
Barbaricum is a rapidly growing government contractor providing leading-edge support to federal customers, focusing on Defense and National Security mission sets. They are seeking an experienced Senior Site Reliability Engineer to enhance the reliability, availability, and operational performance of IT and cloud systems under the MC&FP Outreach and Digital Enterprise Services contract.
Responsibilities:
• Monitor and maintain system reliability, availability, and performance across on-premises, cloud, and hybrid IT environments supporting MC&FP mission requirements.
• Implement proactive performance monitoring, automated alerting, incident response workflows, and resilience engineering practices to reduce downtime and improve operational visibility.
• Develop, maintain, and improve scalable automated infrastructure solutions that support reliable system operations and repeatable service delivery.
• Implement rollback strategies, recovery approaches, and chaos engineering practices to validate resilience, reduce operational risk, and improve system stability.
• Analyze usage patterns, capacity trends, and performance indicators to support dynamic scaling, resource optimization, and system improvement decisions.
• Develop and maintain real-time operational dashboards, reports, and metrics that enable rapid decision-making, leadership awareness, and system optimization.
• Respond to and resolve system outages, impairments, and service disruptions while coordinating with technical teams to minimize mission impact.
• Conduct post-incident reviews to identify root causes, document lessons learned, and implement preventative measures that reduce recurrence.
• Collaborate with software developers, cloud engineers, cybersecurity personnel, and operations teams to improve services, reliability patterns, deployment practices, and operational standards.
• Create and maintain system documentation, configuration standards, operational runbooks, monitoring procedures, and service reliability guidance.
• Automate common operations tasks to reduce manual workloads, improve consistency, and increase system efficiency.
• Implement security best practices across operational activities, infrastructure automation, monitoring, incident response, and system administration functions.
Qualifications:
Required:
• Bachelor's degree in Computer Science, Information Technology, Systems Engineering, Cybersecurity, or a related field.
• 10+ years of experience in site reliability engineering, systems administration, infrastructure operations, cloud operations, DevSecOps, or a similar technical role, particularly in a government, federal, defense, or secure IT setting.
• Expert knowledge of site reliability engineering practices, system monitoring, incident management, automation, performance tuning, and operational resilience.
• Strong understanding of Windows and Linux administration, infrastructure operations, system configuration, service management, and troubleshooting practices.
• Experience with automation platforms and configuration management tools such as Ansible, Puppet, Chef, or similar technologies.
• Proficiency with scripting languages such as Python, Shell, PowerShell, or similar tools used to automate operational and infrastructure tasks.
• Knowledge of cloud services and infrastructure across AWS, Microsoft Azure, Google Cloud, or comparable cloud environments.
• Strong understanding of network troubleshooting, configuration, connectivity analysis, system dependencies, and performance bottleneck identification.
• Ability to design, interpret, and maintain dashboards, alerts, metrics, logs, and operational reporting that support service health and decision-making.
• Ability to conduct root cause analysis, post-incident reviews, and corrective action planning in complex technical environments.
• Strong problem-solving skills and the ability to work under pressure during outages, impairments, and time-sensitive operational issues.
• Excellent written and verbal communication skills, with the ability to explain technical findings, incident impacts, and reliability recommendations to technical and non-technical stakeholders.
• Demonstrated experience maintaining reliable, scalable, and efficiently managed IT systems across on-premises, cloud, or hybrid environments.
• Experience developing automated infrastructure, operational scripts, monitoring solutions, dashboards, runbooks, and configuration standards.
• Experience supporting incident response, system outage resolution, post-incident reviews, root cause analysis, and operational improvement initiatives.
• Experience collaborating with development, infrastructure, cloud, cybersecurity, and program teams to improve reliability, security, and service performance.
• DoD Secret Security Clearance.
Preferred:
• Master's degree preferred.
• Certifications related to cloud computing, system administration, site reliability engineering, DevSecOps, or automation are beneficial.
Company:
Barbaricum is a government relations company that offers strategic communications, research, and analysis solutions. Founded in 2008, the company is headquartered in Washington, USA, with a team of 201-500 employees. The company is currently Growth Stage.