1

Senior Observability Engineer Jobs in Connecticut

The Sr. Digital Software Engineer leads the delivery of complex software solutions while guiding ... Familiarity with DevOps practices, CI/CD pipelines, and observability tools. * Effective ...

Sr. Infrastructure Engineer

Norwalk, CT · On-site

$108K - $148K/yr

Sr. Infrastructure Engineer * This is a Short Term Assignment - 8 - 10 weeks - projected start 8/24 ... Using observability tools, ensure high availability, and quickly respond to incidents and outages.

Sr. Software Engineer

Stamford, CT · On-site

$130K - $172K/yr

As our Senior Software Engineer, you'll join the Content Delivery Engineering team (CDE) within ... Leverage observability tooling (metrics, logs, traces) and performance analytics to detect issues ...

Sr. Software Engineer

Stamford, CT · On-site

$130K - $172K/yr

As our Senior Software Engineer, you'll join the Content Delivery Engineering team (CDE) within ... Leverage observability tooling (metrics, logs, traces) and performance analytics to detect issues ...

Sr Data Engineer

Shelton, CT · On-site

$114K - $137K/yr

The Sr Data Engineer is a senior-level, hands-on technical leader responsible for designing ... DQ, observability, CI/CD). * Drive emerging tech (Iceberg, Lakeflow, Openflow, Cortex, Mosaic AI ...

Be Seen First

Sr. Infrastructure Engineer - Norwalk, CT * This is an 8 - 10 weeks assignment with potential to ... Using observability tools, ensure high availability, and quickly respond to incidents and outages.

Improve platform reliability, scalability, observability, and operational resilience * Establish ... At least 12 years of experience as Staff Engineer or senior-level individual contributor supporting ...

Showing results 21-40

Senior Observability Engineer information

What is a senior observability engineer?

A Senior Observability Engineer is a seasoned IT professional responsible for designing, implementing, and maintaining systems that monitor and provide insights into the performance, health, and reliability of software applications and infrastructure. They utilize tools for logging, monitoring, tracing, and alerting to ensure that systems are observable and any issues can be quickly detected and resolved. In addition to technical expertise, they often collaborate with development and operations teams to establish best practices, improve incident response, and optimize system performance. Their work is crucial for maintaining uptime, enhancing customer experiences, and supporting the scalability of technology platforms.

How does a senior observability engineer typically collaborate with development and operations teams?

A Senior Observability Engineer works closely with both development and operations teams to ensure robust monitoring, logging, and tracing solutions are in place across all applications and infrastructure. They often participate in architecture discussions to advise on best practices for instrumenting code and systems for observability. By analyzing metrics and alerting patterns, they help teams proactively resolve issues and optimize system performance. This role also involves mentoring engineers on observability tools and fostering a culture of transparency and accountability in incident response.

What are the key skills and qualifications needed to thrive as a senior observability engineer, and why are they important?

To thrive as a Senior Observability Engineer, you need expertise in monitoring, logging, and tracing systems, with a solid background in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, and cloud platforms, as well as certifications such as AWS Certified DevOps Engineer, are typically required. Strong problem-solving, collaboration, and communication skills are critical for effectively diagnosing and resolving complex infrastructure issues. These skills ensure reliable system performance, rapid incident response, and continuous improvement of the technology environment.

What is the difference between Senior Observability Engineer vs Site Reliability Engineer?

AspectSenior Observability EngineerSite Reliability Engineer
CredentialsExperience with monitoring tools, scripting, cloud platformsSame as Senior Observability Engineer, often with SRE certifications
Work EnvironmentFocus on monitoring, logging, and tracing systemsFocus on system reliability, automation, and incident response
Industry UsageUsed in tech companies emphasizing system observabilityCommon in large-scale tech and cloud services
Search/Comparison IntentOften compared for monitoring rolesCompared for reliability and system stability roles

While both roles require expertise in cloud platforms and scripting, the Senior Observability Engineer primarily focuses on designing and maintaining monitoring, logging, and tracing systems to ensure system visibility. In contrast, a Site Reliability Engineer emphasizes system reliability, automation, and incident management to maintain service uptime. Both roles are vital in tech environments but serve different core functions related to system health and stability.

How much do senior observability engineers make?

Senior observability engineers typically earn between $110,000 and $160,000 annually, depending on experience, location, and company size. They often work with tools like Prometheus, Grafana, and cloud platforms, and may require advanced knowledge of monitoring, logging, and alerting systems.

What does a senior observability engineer do?

A senior observability engineer designs, implements, and maintains systems to monitor the performance and health of software applications and infrastructure. They utilize tools like Prometheus, Grafana, and ELK stack to analyze metrics, logs, and traces, ensuring system reliability and performance. This role often requires strong scripting skills and knowledge of cloud environments and distributed systems.

What are the most commonly searched types of Observability Engineer jobs in Connecticut?

The most popular types of Observability Engineer jobs in Connecticut are:

What are popular job titles related to Senior Observability Engineer jobs in Connecticut?

For Senior Observability Engineer jobs in Connecticut, the most frequently searched job titles are:

What job categories do people searching Senior Observability Engineer jobs in Connecticut look for?

The top searched job categories for Senior Observability Engineer jobs in Connecticut are:

What cities in Connecticut are hiring for Senior Observability Engineer jobs?

Cities in Connecticut with the most Senior Observability Engineer job openings:

Principal Reliability Engineer - EDS

The Hartford Financial Services Group, Inc.

Hartford, CT • On-site, Remote

Full-time

Re-posted 22 days ago


The Hartford rating

8.8

Company rating: 8.8 out of 10

Based on 121 frontline employees who took The Breakroom Quiz

56th of 311 rated insurance


Job description

Principal Reliability Engineering - IE06JE
We're determined to make a difference and are proud to be an insurance company that goes well beyond coverages and policies. Working here means having every opportunity to achieve your goals - and to help others accomplish theirs, too. Join our team as we help shape the future.
The Enterprise Data Services (EDS) organization is seeking a Principal Reliability Engineer (Principal RE) to serve as the senior technical authority responsible for the reliability, resilience, availability, and performance of all data platforms, cloud infrastructure, data products, and data pipelines across the enterprise data organization. This role sets the strategic vision for Reliability Engineering within EDS and leads the definition, implementation, and continuous evolution of RE practices, tooling, automation, observability frameworks, and AIOps/AI-driven operations.
As the Principal RE, you will influence architectural direction, lead large-scale, cross-organizational technical initiatives, and drive a culture of engineering excellence, automation-first operations, and proactive reliability improvement. You will partner closely with platform engineering, data engineering, security, architecture, and product teams to embed RE principles into every stage of the data product lifecycle.
This role will have a Hybrid work schedule, with the expectation of working in an office (Columbus, OH, Chicago, IL, Hartford, CT or Charlotte, NC) 3 days a week (Tuesday through Thursday).
Key Responsibilities
Enterprise Reliability Strategy & Leadership
  • Work closely with the AVP, RE & Production Support, EDS defining the Reliability Engineering strategy for data platforms, data cloud environments, and data products.
  • Establish long-term RE roadmaps, target operating models, and architectural patterns that scale with organizational growth.
  • Serve as the highest-level technical escalation point for systemic reliability issues, influencing executive stakeholders and engineering leaders.

Platform & Cloud Reliability (AWS, GCP, Snowflake, EMR, Hadoop, ETL/ELT)
  • Leverage Enterprise provided standards and building blocks to Architect and evolve highly reliable, performant, and cost-efficient cloud-based platforms across AWS and GCP for all EDS services.
  • Influence and work directly with Platform Solution Architecture on new product enablement, hyper automation (end to end blueprint automation).
  • Oversee reliability controls and fail-safe patterns for Snowflake, EMR, Hadoop/Spark clusters, container platforms (e.g., Kubernetes), and mission-critical data systems.
  • Lead the creation and enforcement of SLO/SLI frameworks that span the entire data lifecycle.

AI-Enabled Operations, AIOps & Intelligent Automation
  • Develop and implement AI-driven automation for anomaly detection, alert correlation, autonomous remediation, and predictive capacity management.
  • Leverage LLMs, prompt engineering, and cloud-native AI services (AWS Bedrock, SageMaker, Vertex AI) to build intelligent runbooks, advanced troubleshooting agents, and generative-AI-enabled operational tooling.
  • Champion the adoption of machine learning-based observability and reliability analytics.

End-to-End Observability & Operational Excellence
  • Adopt and architect enterprise-wide data observability frameworks-including logging, metrics, tracing, distributed profiling, and event pipelines-for all data platforms and pipelines.
  • Establish gold-standard incident response patterns, post-incident reviews, and continuous improvement processes.
  • Drive elimination of toil across EDS, focusing on self-healing systems, proactive detection, and autonomous operations.

Data Pipeline & Data Product Reliability
  • Define RE best practices for modern data products, governed data pipelines, real-time/streaming systems, and operational analytics platforms.
  • Ensure data quality, data timeliness, and SLAs for data products through automated checks, lineage-informed alerting, and pipeline reliability tooling.
  • Partner with Data Engineering to embed resilience patterns (idempotency, checkpointing, replayability, disaster recovery) into pipeline architectures.

Engineering Standards, Governance & Cross-Org Influence
  • Set and enforce standards for IaC, CI/CD, platform automation, reliability frameworks, operational readiness, and runbook quality across EDS.
  • Provide technical leadership and mentorship to Staff/Senior Engineers in the RE team and Production Support teams, influencing engineering culture and helping grow RE capabilities across the organization.
  • Represent Reliability Engineering in architectural reviews, enterprise governance forums, and executive-level discussions.

Technical Experience
  • 10+ years in one or more of the following areas: data, cloud, platform engineering, site/reliability engineering, or large-scale distributed systems, with experience in leadership or technology leader roles.
  • Proficiency with data or cloud platforms, including architectural patterns for resilience, networking, security, and distributed data infrastructure.
  • Deep experience supporting or engineering platforms such as Snowflake, EMR, Hadoop/Spark, Data Integration, and cloud-native data ecosystems.
  • Scripting and programming (preferably Python) for large-scale automation, platform tooling, and reliability frameworks.
  • Experience with Infrastructure-as-Code (Terraform, CloudFormation) and enterprise CI/CD.

Preferred Qualifications
  • Experience in regulated or highly complex enterprise environments (financial services, insurance, healthcare).
  • Prior experience as a Senior Staff Engineer, Engineering or Architecture leader with hands on experience, or similar senior technical role.
  • Knowledge of data governance, metadata, lineage systems, and data quality engineering practices.
  • Certifications in AWS, GCP, Kubernetes, or SRE/DevOps frameworks.

AI & AIOps
  • Background applying machine learning to operations-anomaly detection, event correlation, predictive modeling, and automated remediation.
  • Understand of AI-enabled developer/operations tools using LLMs, prompt engineering, or cloud AI services for reliability improvements.

Observability & Platform Operations
  • Expertise with enterprise observability stacks (Prometheus, Grafana, Datadog, Splunk, Dynatrace, OpenTelemetry).
  • Ability to design and enforce advanced SLI/SLO frameworks across complex data ecosystems.

Leadership & Cross-Functional Influence
  • Demonstrated ability to lead technical strategy at scale, influence senior engineering leaders, and set enterprise-wide standards.
  • Strong capability in mentoring engineers, providing architectural guidance, and fostering engineering excellence.
  • Exceptional communication skills for interacting with executives, senior architects, product leaders, and engineering teams.

Candidate must be authorized to work in the US without company sponsorship. The company will not support the STEM OPT I-983 Training Plan endorsement for this position.
Compensation
The listed annualized base pay range is primarily based on analysis of similar positions in the external market. Actual base pay could vary and may be above or below the listed range based on factors including but not limited to performance, proficiency and demonstration of competencies required for the role. The base pay is just one component of The Hartford's total compensation package for employees. Other rewards may include short-term or annual bonuses, long-term incentives, and on-the-spot recognition. The annualized base pay range for this role is:
$152,800 - $229,200
Equal Opportunity Employer/Sex/Race/Color/Veterans/Disability/Sexual Orientation/Gender Identity or Expression/Religion/Age
About Us | Our Culture | What It's Like to Work Here | Perks & Benefits

What The Hartford employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Hartford logo

About Hartford

Sourced by ZipRecruiter

Hartford Financial Services Group, widely recognized as The Hartford, is a renowned company based in Hartford, CT, US. Established in 1810, it has evolved into an industry leader in the insurance and financial services sector, proudly serving more than one million businesses in the US. The Hartford is committed to offering a gamut of insurance products that include homeowners, automobile, and business insurance as well as employee benefits and mutual funds. The company’s core values revolve around customer-focused innovations, diversity and inclusion, and ethical dealings that have earned them a customer-centric reputation. This shapes their mission which revolves around aiding their clients to overcome unforeseen obstacles and enhancing their wealth over time. Among the company's noted accomplishments is being consistently listed among the World's Most Ethical Companies, a testament to their unwavering commitment towards responsible business practices.

Industry

Finance and insurance

Company size

10,000+ Employees

Headquarters location

Hartford, CT, US

Year founded

1810

Social media