... management. * Leverage LLMs, prompt engineering, and cloudnative AI services (AWS Bedrock ... Certifications in AWS, GCP, Kubernetes, or SRE/DevOps frameworks. AI & AIOps * Background applying ...
... management. * Leverage LLMs, prompt engineering, and cloudnative AI services (AWS Bedrock ... Certifications in AWS, GCP, Kubernetes, or SRE/DevOps frameworks. AI & AIOps * Background applying ...
... (SRE) best practices to maximize uptime and quickly resolve incidents.Drive continuous improvementin operational processes (change management, incident response) to enhance system reliability ...
... (SRE) best practices to maximize uptime and quickly resolve incidents.Drive continuous improvementin operational processes (change management, incident response) to enhance system reliability ...
... (SRE) best practices to maximize uptime and quickly resolve incidents.Drive continuous improvementin operational processes (change management, incident response) to enhance system reliability ...
... (SRE) best practices to maximize uptime and quickly resolve incidents.Drive continuous improvementin operational processes (change management, incident response) to enhance system reliability ...
Software Engineer
Stamford, CT · Hybrid
Working closely with our SRE team to ensure deployed systems are reliable, resilient, scalable, and perform optimally, using common, modern technologies. * Make best use of our observability systems ...
Quick apply
Software Engineer
Stamford, CT · Hybrid
Working closely with our SRE team to ensure deployed systems are reliable, resilient, scalable, and perform optimally, using common, modern technologies. * Make best use of our observability systems ...
Software Engineer
Stamford, CT · Hybrid
Working closely with our SRE team to ensure deployed systems are reliable, resilient, scalable, and perform optimally, using common, modern technologies. * Make best use of our observability systems ...
Software Engineer
Stamford, CT · Hybrid
Working closely with our SRE team to ensure deployed systems are reliable, resilient, scalable, and perform optimally, using common, modern technologies. * Make best use of our observability systems ...
Mission Assurance Engineer - Survivability & Reliability
Danbury, CT · On-site
$104K - $132K/yr
The Analyst function supports the Specialty Engineering Manager in Survivability/Reliability ... This position is an on-site position located in Danbury, CT. Responsibilities: * Reviewing Customer ...
Mission Assurance Engineer - Survivability & Reliability
Danbury, CT · On-site
$104K - $132K/yr
The Analyst function supports the Specialty Engineering Manager in Survivability/Reliability ... This position is an on-site position located in Danbury, CT. Responsibilities: * Reviewing Customer ...
Partner with SRE/DevOps to define reliability-aligned targets and run controlled chaos-engineering experiments (fault/latency injection, dependency degradation) to validate resilience and recovery;
New
Partner with SRE/DevOps to define reliability-aligned targets and run controlled chaos-engineering experiments (fault/latency injection, dependency degradation) to validate resilience and recovery;
New
Co-op- Reliability Engineer Fall 2025 (Internship)
Wilton, CT · On-site
$106K - $133K/yr
... growth management. As a Co-Op in the Quality, Reliability and Testing Department, you will ... This position is located on-site in Wilton, Connecticut. It requires onsite presence to attend in ...
Co-op- Reliability Engineer Fall 2025 (Internship)
Wilton, CT · On-site
$106K - $133K/yr
... growth management. As a Co-Op in the Quality, Reliability and Testing Department, you will ... This position is located on-site in Wilton, Connecticut. It requires onsite presence to attend in ...
Co-op- Reliability Engineer Fall 2025 (Internship)
Wilton, CT · On-site
$106K - $133K/yr
... growth management. As a Co-Op in the Quality, Reliability and Testing Department, you will ... This position is located on-site in Wilton, Connecticut. It requires onsite presence to attend in ...
Co-op- Reliability Engineer Fall 2025 (Internship)
Wilton, CT · On-site
$106K - $133K/yr
... growth management. As a Co-Op in the Quality, Reliability and Testing Department, you will ... This position is located on-site in Wilton, Connecticut. It requires onsite presence to attend in ...
Chief Technology, Infrastructure & Product Engineering Officer
Stamford, CT · On-site
$275 - $300/hr
Ensure availability, performance and resilience of critical services through SRE practices. * Drive proactive monitoring, observability, capacity management and automated remediation. * Improve ...
New
Chief Technology, Infrastructure & Product Engineering Officer
Stamford, CT · On-site
$275 - $300/hr
Ensure availability, performance and resilience of critical services through SRE practices. * Drive proactive monitoring, observability, capacity management and automated remediation. * Improve ...
New
Chief Technology, Infrastructure & Product Engineering Officer
Hartford, CT · On-site
$275 - $300/hr
Ensure availability, performance and resilience of critical services through SRE practices. * Drive proactive monitoring, observability, capacity management and automated remediation. * Improve ...
New
Chief Technology, Infrastructure & Product Engineering Officer
Hartford, CT · On-site
$275 - $300/hr
Ensure availability, performance and resilience of critical services through SRE practices. * Drive proactive monitoring, observability, capacity management and automated remediation. * Improve ...
New
Summary The Reliability Engineer will work closely with Sibelco plant maintenance, production and ... reliability of site. * Track the performance of reliability improvements to determine their ...
Quick apply
Summary The Reliability Engineer will work closely with Sibelco plant maintenance, production and ... reliability of site. * Track the performance of reliability improvements to determine their ...
Splunk Developer
Hartford, CT · On-site
... or Site Reliability Engineering roles.Hands-on experience with Splunk for log analysis and monitoring.Strong understanding of log management and analysis techniques.Solid programming or scripting ...
Quick apply
Splunk Developer
Hartford, CT · On-site
... or Site Reliability Engineering roles.Hands-on experience with Splunk for log analysis and monitoring.Strong understanding of log management and analysis techniques.Solid programming or scripting ...
Institutionalize SRE fundamentals and advance Infrastructure as Code and GitOps. * Embed security and privacy by design-partner with Security/Risk to implement Zero Trust, secrets and key management ...
Institutionalize SRE fundamentals and advance Infrastructure as Code and GitOps. * Embed security and privacy by design-partner with Security/Risk to implement Zero Trust, secrets and key management ...
Institutionalize SRE fundamentals and advance Infrastructure as Code and GitOps. * Embed security and privacy by design-partner with Security/Risk to implement Zero Trust, secrets and key management ...
Institutionalize SRE fundamentals and advance Infrastructure as Code and GitOps. * Embed security and privacy by design-partner with Security/Risk to implement Zero Trust, secrets and key management ...
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
Senior Mission Assurance Engineer - Survivability, Reliability & Components {D}
Danbury, CT · On-site
This position is an on-site position located in Danbury, CT. We offer generous relocation benefits ... Managing the Preliminary Electrical Parts List and supporting the Supply Chain procurement efforts.
We value a "manager of one" mindset, where outcomes matter more than optics. Authority is earned ... Site Reliability Engineering (SRE) * Quality engineering / testing strategy * Secure SDLC and ...
We value a "manager of one" mindset, where outcomes matter more than optics. Authority is earned ... Site Reliability Engineering (SRE) * Quality engineering / testing strategy * Secure SDLC and ...
Site Reliability Engineer Manager information
See Connecticut salary details
$10.29 - $17.30
1% of jobs
$17.30 - $24.30
0% of jobs
$24.30 - $31.31
0% of jobs
$31.31 - $38.31
2% of jobs
$38.31 - $45.32
4% of jobs
$45.32 - $52.32
17% of jobs
$52.39 is the 25th percentile. Wages below this are outliers.
$52.32 - $59.33
27% of jobs
$59.33 - $66.34
20% of jobs
$67.54 is the 75th percentile. Wages above this are outliers.
$66.34 - $73.34
17% of jobs
$73.34 - $80.35
6% of jobs
$80.35 - $87.35
4% of jobs
$10
$60
$87
How much do site reliability engineer manager jobs pay per hour?
What is a site reliability engineer manager?
What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?
| Aspect | Site Reliability Engineer (SRE) | Site Reliability Engineer Manager |
|---|---|---|
| Responsibilities | Focuses on designing, implementing, and maintaining reliable systems and automation | Oversees SRE teams, manages projects, and aligns reliability goals with business objectives |
| Required Skills | Strong coding, system design, and troubleshooting skills | Leadership, team management, strategic planning |
| Certifications | Google Cloud, AWS certifications, Linux, scripting | Same as SRE, plus management certifications (e.g., PMP) often preferred |
| Work Environment | Technical, hands-on with systems and automation | Managerial, coordinating teams and projects |
The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.
How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?
What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

The Hartford rating
8.8
Based on 121 frontline employees who took The Breakroom Quiz
57th of 304 rated insurance
Job description
We're determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals - and to help others accomplish theirs, too. Join our team as we help shape the future.
The Enterprise Data Services (EDS) organization is seeking a Principal Reliability Engineer (Principal RE) to serve as the senior technical authority responsible for the reliability, resilience, availability, and performance of all data platforms, cloud infrastructure, data products, and data pipelines across the enterprise data organization. This role sets the strategic vision for Reliability Engineering within EDS and leads the definition, implementation, and continuous evolution of RE practices, tooling, automation, observability frameworks, and AIOps/AIdriven operations.
As the Principal RE, you will influence architectural direction, lead largescale, crossorganizational technical initiatives, and drive a culture of engineering excellence, automationfirst operations, and proactive reliability improvement. You will partner closely with platform engineering, data engineering, security, architecture, and product teams to embed RE principles into every stage of the data product lifecycle.
This role will have a Hybrid work schedule, with the expectation of working in an office (Columbus, OH, Chicago, IL, Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).
Key Responsibilities
Enterprise Reliability Strategy & Leadership
- Work closely with the AVP, RE & Production Support, EDS defining the Reliability Engineering strategy for data platforms, data cloud environments, and data products.
- Establish longterm RE roadmaps, target operating models, and architectural patterns that scale with organizational growth.
- Serve as the highestlevel technical escalation point for systemic reliability issues, influencing executive stakeholders and engineering leaders.
Platform & Cloud Reliability (AWS, GCP, Snowflake, EMR, Hadoop, ETL/ELT)
- Leverage Enterprise provided standards and building blocks to Architect and evolve highly reliable, performant, and costefficient cloudbased platforms across AWS and GCP for all EDS services.
- Influence and work directly with Platform Solution Architecture on new product enablement, hyper automation (end to end blueprint automation).
- Oversee reliability controls and failsafe patterns for Snowflake, EMR, Hadoop/Spark clusters, container platforms (e.g., Kubernetes), and missioncritical data systems.
- Lead the creation and enforcement of SLO/SLI frameworks that span the entire data lifecycle.
AIEnabled Operations, AIOps & Intelligent Automation
- Develop and implement AIdriven automation for anomaly detection, alert correlation, autonomous remediation, and predictive capacity management.
- Leverage LLMs, prompt engineering, and cloudnative AI services (AWS Bedrock, SageMaker, Vertex AI) to build intelligent runbooks, advanced troubleshooting agents, and generativeAIenabled operational tooling.
- Champion the adoption of machine learning-based observability and reliability analytics.
EndtoEnd Observability & Operational Excellence
- Adopt and architect enterprisewide data observability frameworks-including logging, metrics, tracing, distributed profiling, and event pipelines-for all data platforms and pipelines.
- Establish goldstandard incident response patterns, postincident reviews, and continuous improvement processes.
- Drive elimination of toil across EDS, focusing on selfhealing systems, proactive detection, and autonomous operations.
Data Pipeline & Data Product Reliability
- Define RE best practices for modern data products, governed data pipelines, realtime/streaming systems, and operational analytics platforms.
- Ensure data quality, data timeliness, and SLAs for data products through automated checks, lineage-informed alerting, and pipeline reliability tooling.
- Partner with Data Engineering to embed resilience patterns (idempotency, checkpointing, replayability, disaster recovery) into pipeline architectures.
Engineering Standards, Governance & CrossOrg Influence
- Set and enforce standards for IaC, CI/CD, platform automation, reliability frameworks, operational readiness, and runbook quality across EDS.
- Provide technical leadership and mentorship to Staff/Senior Engineers in the RE team and Production Support teams, influencing engineering culture and helping grow RE capabilities across the organization.
- Represent Reliability Engineering in architectural reviews, enterprise governance forums, and executivelevel discussions.
Technical Experience
- 10+ years in one or more of the following areas: data, cloud, platform engineering, site/reliability engineering, or largescale distributed systems, with experience in leadership or technology leader roles.
- Proficiency with data or cloud platforms, including architectural patterns for resilience, networking, security, and distributed data infrastructure.
- Deep experience supporting or engineering platforms such as Snowflake, EMR, Hadoop/Spark, Data Integration, and cloudnative data ecosystems.
- Scripting and programming (preferably Python) for largescale automation, platform tooling, and reliability frameworks.
- Experience with InfrastructureasCode (Terraform, CloudFormation) and enterprise CI/CD.
Preferred Qualifications
- Experience in regulated or highly complex enterprise environments (financial services, insurance, healthcare).
- Prior experience as a Senior Staff Engineer, Engineering or Architecture leader with hands on experience, or similar senior technical role.
- Knowledge of data governance, metadata, lineage systems, and data quality engineering practices.
- Certifications in AWS, GCP, Kubernetes, or SRE/DevOps frameworks.
AI & AIOps
- Background applying machine learning to operations-anomaly detection, event correlation, predictive modeling, and automated remediation.
- Understand of AIenabled developer/operations tools using LLMs, prompt engineering, or cloud AI services for reliability improvements.
Observability & Platform Operations
- Expertise with enterprise observability stacks (Prometheus, Grafana, Datadog, Splunk, Dynatrace, OpenTelemetry).
- Ability to design and enforce advanced SLI/SLO frameworks across complex data ecosystems.
Leadership & CrossFunctional Influence
- Demonstrated ability to lead technical strategy at scale, influence senior engineering leaders, and set enterprisewide standards.
- Strong capability in mentoring engineers, providing architectural guidance, and fostering engineering excellence.
- Exceptional communication skills for interacting with executives, senior architects, product leaders, and engineering teams.
Candidate must be authorized to work in the US without company sponsorship.The company will not support the STEM OPT I-983 Training Plan endorsement for this position.
Compensation
The listed annualized base pay range is primarily based on analysis of similar positions in the external market. Actual base pay could vary and may be above or below the listed range based on factors including but not limited to performance, proficiency and demonstration of competencies required for the role. The base pay is just one component of The Hartford's total compensation package for employees. Other rewards may include short-term or annual bonuses, long-term incentives, and on-the-spot recognition. The annualized base pay range for this role is:
$152,800 - $229,200Equal Opportunity Employer/Sex/Race/Color/Veterans/Disability/Sexual Orientation/Gender Identity or Expression/Religion/Age
About Us|Our Culture|What It's Like to Work Here|Perks & Benefits
What The Hartford employees say
Pay
Benefits
Hours and flexibility
Workplace
Get the full story on Breakroom
About Hartford
Sourced by ZipRecruiter
Hartford Financial Services Group, widely recognized as The Hartford, is a renowned company based in Hartford, CT, US. Established in 1810, it has evolved into an industry leader in the insurance and financial services sector, proudly serving more than one million businesses in the US. The Hartford is committed to offering a gamut of insurance products that include homeowners, automobile, and business insurance as well as employee benefits and mutual funds. The company’s core values revolve around customer-focused innovations, diversity and inclusion, and ethical dealings that have earned them a customer-centric reputation. This shapes their mission which revolves around aiding their clients to overcome unforeseen obstacles and enhancing their wealth over time. Among the company's noted accomplishments is being consistently listed among the World's Most Ethical Companies, a testament to their unwavering commitment towards responsible business practices.
Industry
Finance and insurance
Company size
10,000+ Employees
Headquarters location
Hartford, CT, US
Year founded
1810