1

Observability Sre Jobs (NOW HIRING)

Site Reliability Engineer

Buffalo, NY ยท On-site

$100 - $110/hr

The ideal candidate will have strong experience in observability, automation, incident management, Azure, and Infrastructure as Code and a proven ability to design, implement, and mature SRE ...

Site Reliability Engineer

Plano, TX ยท On-site

$54.50 - $72.50/hr

Site Reliability Engineer Hybrid 3 times a week in Iselin, NJ OR Hybrid 3 times a week in PLANO, TX ... Implement observability, logging, and monitoring using tools like Prometheus, Grafana, ELK, or ...

Site Reliability Engineer

Beaverton, OR ยท Hybrid

$59.25 - $78.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Overview As our Site Reliability Engineer, you'll help drive Concora Credit's Mission to enable ... Contribute to cloud reliability through automation, observability, incident reduction, capacity ...

Showing results 41-60

Observability Sre information

See salary details

$10

$63

$91

How much do observability sre jobs pay per hour?

As of Aug 14, 2026, the average hourly pay for observability sre in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

What is the difference between Observability Sre vs Site Reliability Engineer?

AspectObservability SreSite Reliability Engineer
Primary FocusMonitoring, logging, and tracing to ensure system observabilitySystem reliability, automation, and infrastructure management
Skills & CertificationsMonitoring tools, scripting, cloud platforms, observability frameworksLinux, scripting, cloud services, automation tools
Work EnvironmentCollaborates with SRE, DevOps, and development teams on observability practicesBuilds and maintains scalable, reliable systems in production

While both roles focus on system stability, Observability Sre specializes in monitoring and diagnostics, whereas Site Reliability Engineers focus on overall system reliability and automation. They often work together to ensure robust, observable, and reliable systems.

What is an Observability SRE?

An Observability SRE (Site Reliability Engineer) is a specialist focused on ensuring that systems and applications are transparent, measurable, and reliable. Their main responsibility is to implement and maintain tools for monitoring, logging, and tracing, providing insights into system performance and health. Observability SREs help teams quickly detect, diagnose, and resolve issues by making system behavior visible and understandable. They play a critical role in uptime, incident response, and performance optimization, bridging the gap between software development and IT operations.

What are some typical challenges faced by Observability SREs when implementing monitoring solutions across diverse systems?

Observability SREs often encounter challenges when integrating monitoring tools across varied technology stacks and legacy systems. Ensuring consistent data collection, standardizing metrics, and maintaining visibility in complex, distributed environments can be difficult. Collaborating with development and operations teams to define meaningful alerts and dashboards requires strong communication and a deep understanding of both infrastructure and application behaviors. Staying up-to-date with evolving tools and best practices is also essential to address emerging observability needs.

What are the key skills and qualifications needed to thrive as an Observability SRE?

To thrive as an Observability SRE, you need a solid background in systems engineering, monitoring best practices, and expertise in observability concepts, often supported by a degree in computer science or related fields. Familiarity with tools like Prometheus, Grafana, ELK stack, and cloud monitoring platforms, as well as scripting languages such as Python or Bash, is typically required. Strong problem-solving, collaboration, and communication skills help SREs respond to incidents and work across teams effectively. These skills ensure system reliability, rapid issue detection, and continuous service improvement in complex technical environments.
More about Observability Sre jobs

What cities are hiring for Observability Sre jobs?

Cities with the most Observability Sre job openings:

What states have the most Observability Sre jobs?

States with the most job openings for Observability Sre jobs include:

Infographic showing various Observability Sre job openings in the United States as of August 2026, with employment types broken down into 60% Full Time, and 40% Contract. Highlights an 60% In-person, and 40% Remote job distribution, with an average salary of $132,583 per year, or $63.7 per hour.

Site Reliability Engineer

BC Forward

Buffalo, NY โ€ข On-site

$100 - $110/hr

Other

Posted 19 days ago


Job description

Job Title: Site Reliability Engineer Location: Buffalo NY Duration: Temp - 12 months Pay Range: $100/hr $110/hr (W2) Job ID: 407436 About BCforward BCforward is a leading global IT consulting and workforce solutions firm providing services and support to Fortune 500 and government clients. Founded in 1998, BCforward has grown with our customers needs into a full-service business solutions provider. With delivery centers and offices across North America and India, we take pride in building long-term relationships and delivering excellence through innovation, collaboration, and integrity. Job Description We are seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, performance, and operational excellence of critical banking platforms and applications. The ideal candidate will have strong experience in observability, automation, incident management, Azure, and Infrastructure as Code and a proven ability to design, implement, and mature SRE practices across the SDLC while leading complex reliability initiatives. Responsibilities: Design, implement, and support highly available, scalable, and resilient applications and cloud infrastructure aligned to enterprise standards and SRE best practices. Lead initiatives to improve reliability, availability, performance, and operational maturity through automation and engineering excellence. Define, implement, and monitor SLOs, SLIs, and error budgets for critical services. Develop observability strategies using Dynatrace, OpenTelemetry, distributed tracing, metrics, logs, dashboards, and alerting. Design and maintain end-to-end monitoring that provides actionable insights into application, infrastructure, and customer experience health. Analyze production telemetry to identify performance bottlenecks, reliability risks, and capacity constraints proactively. Lead incident response for high-severity events and coordinate cross-functional restoration and communications. Perform and facilitate RCAs with corrective and preventive actions tracked to completion. Automate repetitive tasks, operational workflows, deployments, recovery procedures, and reliability controls. Partner with development teams to embed reliability and observability across the SDLC. Design, develop, and execute automated regression testing to validate stability and performance after changes. Review test coverage and reliability validation to ensure comprehensive risk mitigation. Create, maintain, and improve Terraform-based IaC for provisioning, configuration, and standardization. Support and optimize Microsoft Azure environments, including App Services, resource management, scaling, and deployment automation. Use Azure Monitor, Application Insights, and Log Analytics to improve visibility and reliability. Drive performance testing, resiliency testing, fault tolerance validation, and disaster recovery preparedness. Establish operational readiness standards and enforce requirements before production deployments. Review architectures and recommend improvements for resiliency, efficiency, and cloud optimization. Lead capacity planning, performance tuning, and workload optimization across production environments. Develop and maintain runbooks, incident playbooks, knowledge articles, and SOPs. Partner with engineering, infrastructure, cybersecurity, architecture, and support teams on cross-functional improvements. Communicate system health, reliability trends, risks, and remediation to technical and business stakeholders. Present initiatives, metrics, and recommendations at reviews, forums, and leadership meetings. Mentor engineers on observability, cloud engineering, automation, SRE principles, and operational practices. Adhere to risk and regulatory standards and identify issues requiring escalation. Promote a culture of belonging consistent with company values and maintain internal control standards. Required Skills & Qualifications: Associate's degree with 7+ years in SRE, Cloud, Systems, Infrastructure Engineering, DevOps, or Application Support. Bachelor's degree with 5+ years. Or 9+ years combined education and experience with 5+ years in a technology engineering role. Hands-on observability and monitoring experience with Dynatrace, OpenTelemetry, distributed tracing, metrics, centralized logging, alerting, and dashboards. Proven ability to design and execute automated regression testing frameworks and suites. Strong proficiency with Terraform and Infrastructure as Code practices. Experience with CI/CD, deployment automation, and operational tooling. Expertise in production monitoring, incident management, and troubleshooting of distributed systems. Understanding of application performance management and modern cloud-native architectures. Preferred Skills: Microsoft Azure expertise, including App Services, Resource Groups, networking, scaling and optimization, deployment and release management, and application lifecycle management. Use of Azure Monitor, Application Insights, Log Analytics, dashboards, and alerting. SRE practices such as SLOs, SLIs, error budgets, incident/problem management, RCA, and reliability automation. Experience with performance tuning, capacity planning, proactive issue detection, and observability-driven improvements. Automated recovery mechanisms and self-healing solutions, resiliency patterns, DR planning, and high-availability architectures. Why BCforward? At BCforward, we believe in advancing lives and careers. When you join our team, you gain access to: Competitive compensation and benefits. Opportunities for growth with global clients. A supportive, inclusive culture that values innovation and people. Exposure to modern technologies and projects. About Our Commitment BCforward is an equal opportunity employer. We value diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, gender identity, national origin, age, disability, or veteran status. Interested? Apply Now! If this sounds like the right opportunity for you, please apply with your most recent resume.