1

Observability Site Reliability Engineer Jobs in Florida

Sr. Site Reliability Engineer

Orlando, FL

$53.25 - $70.75/hr

Build observability solutions and dashboards that provide visibility into CI/CD health, deployment success, build performance, developer wait times, and platform reliability. * Partner with ...

Site Reliability Engineer

Tampa, FL

$53.75 - $71.50/hr

\n \n \n Looking for an SRE experienced in monitoring and observability using Prometheus, Grafana, ELK Stack, DataDog, and Sentry. You will assist in converting clients off of Datadog to an Otel ...

SRE Sr Leader- REMOTE

Jacksonville, FL · On-site +1

$75 - $95/hr

SRE Sr Leader Remote 6 Months We are seeking an SRE Senior Leader to drive system uptime ... and enhancing observability. • Incident Management: Drive root cause analysis (RCA), lead ...

Site Reliability Engineer

Deerfield Beach, FL · On-site

$55.25 - $73.25/hr

Site Reliability Engineer Location: Deerfield IL (Onsite) Duration: Full-time only Must Have ... Experience designing and operating enterprise observability platforms using Azure Monitor, Log ...

SRE Lead

Miami, FL · On-site

$54.50 - $72.25/hr

Tata Consultancy Services is seeking an SRE Lead who will provide strategic and technical ... modern observability and cloud infrastructure. Preferred : • Familiarity with modern AIOps ...

Staff Site Reliability Engineer

Boca Raton, FL · On-site

$54 - $72/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog ...

Site Reliability Engineer II

Orlando, FL · On-site

$53.25 - $70.75/hr

Site Reliability Engineer II The SRE II sits at the intersection of software engineering and ... Reliability & Observability * Define, instrument, and enforce SLIs and SLOs in partnership with ...

Own and improve monitoring, alerting, and observability using Grafana , Pingdom , and Uptrends ... Mentor other engineers and help set SRE standards and best practices Required Qualifications * 5+ ...

Own and improve monitoring, alerting, and observability using Grafana , Pingdom , and Uptrends ... Mentor other engineers and help set SRE standards and best practices Required Qualifications * 5+ ...

next page

Showing results 1-20

Observability Site Reliability Engineer information

What engineer makes $500,000 a year?

A senior or principal Site Reliability Engineer (SRE) or Observability Engineer with extensive experience, specialized skills, and working at large tech companies can earn $500,000 or more annually. Compensation often includes base salary, bonuses, and stock options, especially in high-demand markets and organizations with complex infrastructure.

Is AI replacing SRE?

AI is augmenting the work of Site Reliability Engineers (SREs) by automating tasks such as monitoring, incident detection, and response. However, SREs are still essential for designing systems, managing complex issues, and making strategic decisions that require human judgment. AI tools are considered complementary rather than replacements for SREs' expertise and problem-solving skills.

What engineers make $200,000 a year?

Senior Site Reliability Engineers and Observability Engineers with extensive experience, advanced skills in cloud platforms, automation, and monitoring tools can earn $200,000 or more annually. High compensation often correlates with working at large tech companies, possessing specialized certifications, and managing complex, scalable systems.

What is the difference between Observability Site Reliability Engineer vs Monitoring Engineer?

AspectObservability Site Reliability EngineerMonitoring Engineer
FocusEnsuring system reliability through observability, automation, and incident responseImplementing and managing monitoring tools and dashboards
SkillsCloud platforms, scripting, incident management, observability toolsMonitoring tools, alerting systems, data analysis
Work EnvironmentDevOps teams, cloud infrastructure, large-scale systemsOperations teams, infrastructure monitoring

While both roles involve system health, the Observability Site Reliability Engineer focuses on comprehensive system reliability using observability practices, whereas Monitoring Engineers primarily manage monitoring tools and alerts. The SRE role emphasizes automation, incident response, and system resilience, making it broader in scope.

What engineers make $300,000 a year?

Senior Site Reliability Engineers and Observability Engineers with extensive experience, advanced skills in cloud platforms, automation, and monitoring tools can earn $300,000 or more annually. High compensation often correlates with working at large tech companies, possessing specialized certifications, and taking on leadership or highly technical roles.
What job categories do people searching Observability Site Reliability Engineer jobs in Florida look for? The top searched job categories for Observability Site Reliability Engineer jobs in Florida are:
What cities in Florida are hiring for Observability Site Reliability Engineer jobs? Cities in Florida with the most Observability Site Reliability Engineer job openings:

Senior Site Reliability Engineer (SRE) - Observability & Resilience | Hybrid | Full-Time | Glendale,

SKANDA SOLUTIONS LLC

Florida City, FL • On-site

$52.50 - $69.75/hr

Other

Posted 17 days ago


Job description

If you''re interested, please share your updated resume along with:

Current Location

Work Authorization

Expected Yearly full time Salary ?

JOB TITLE : Senior SRE

SKILL CATEGORY : Cloud: AWS

REQUIRED SKILLS : Site Reliability Engineering (SRE) & Kubernetes Operations

WORK LOCATION : (Glendale, Orlando, Seattle)- 3 Days Hybrid

Overview / Summary

We are seeking a Site Reliability Engineer (SRE) with 8–10 years of experience to drive reliability, observability, and resilience improvements across critical systems. This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, continuous automation, and reliability engineering for Tier-1 business-critical applications.

Key Responsibilities

Drive automation initiatives to improve system performance and operational efficiency.

Improve application reliability and availability by proactively identifying and mitigating risks.

Analyze production incidents and root cause analyses (RCAs) to eliminate recurring issues and reduce outages.

Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets using Nobl9.

Conduct reliability assessments across applications, infrastructure, Kubernetes, databases, networks, caching platforms, and cloud environments.

Drive observability improvements using OpenTelemetry, Grafana Cloud, AppDynamics, Splunk, and monitoring best practices.

Perform performance and scalability reviews to support current and future demand.

Lead chaos engineering exercises using Gremlin or Harness Chaos Engineering.

Review cloud architectures against AWS Well-Architected Framework standards and drive remediation of reliability gaps.

Automate operational tasks and implement self-healing solutions.

Identify and eliminate single points of failure (SPOFs) and strengthen disaster recovery and failover capabilities.

Collaborate with Development, Infrastructure, Performance Engineering, and Operations teams to improve system resilience.

Establish reliability governance, dashboards, runbooks, and continuous improvement processes.

Reliability Assessment & Engineering

Conduct application reliability assessments using established reliability frameworks.

Review historical incidents, Sev-1/Sev-2 RCAs, and recurring failure patterns.

Identify reliability debt and drive remediation initiatives.

Evaluate application readiness for SRE engagement.

Perform end-to-end reliability reviews across application, infrastructure, network, and platform layers.

Define reliability roadmaps and track improvement initiatives.

Incident Management & RCA

Analyze incident trends using CSI or equivalent incident management platforms.

Participate in Major Incident Management and Problem Management processes.

Drive RCA reviews and corrective actions.

Track reliability improvement initiatives resulting from postmortems.

Reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

 

Service Level Management

Define and implement SLIs.

Establish SLOs and Error Budgets using Nobl9.

Partner with Product and Engineering teams to define business-focused reliability targets.

Build SLO dashboards and reliability scorecards.

Monitor error budget consumption and enforce governance policies.

Conduct reliability reviews based on SLO compliance.

Cloud & Platform Reliability

Review cloud architectures against AWS Well-Architected Framework principles.

Conduct reliability, performance, cost optimization, security, and operational excellence assessments.

Identify High Risk Issues (HRIs) and drive remediation.

Validate high availability, disaster recovery, backup, and failover capabilities.

Ensure multi-AZ and multi-region deployment strategies are implemented where required.

Kubernetes & Infrastructure Reliability

Review Kubernetes cluster health and workload configurations.

Validate resource requests, limits, autoscaling, and resiliency patterns.

Assess readiness, liveness, and startup probes.

Review service mesh configurations, network policies, and traffic routing.

Validate database high availability, caching strategies, and scaling configurations.

Identify and eliminate single points of failure.

Observability & Monitoring

Design and improve enterprise observability strategies.

Implement OpenTelemetry-based telemetry collection.

Manage metrics, events, logs, and traces (MELT).

Integrate telemetry into Grafana Cloud, Splunk Observability, or equivalent platforms.

Utilize AI-driven observability capabilities for anomaly detection and root cause analysis.

Improve alert quality, reduce alert fatigue, and increase actionable monitoring coverage.

Ensure every alert has an owner, runbook, and customer impact justification.

Application Performance Engineering

Conduct dependency mapping and architecture reviews.

Analyze latency, throughput, and scalability bottlenecks.

Review timeout, retry, circuit breaker, and resilience patterns.

Collaborate with Performance Engineering teams on load and stress testing.

Validate system capacity against current and future traffic demands.

Review Akamai CDN configurations, traffic routing, caching, and failover strategies.

Ensure applications can sustain significant traffic spikes and peak loads.

Chaos Engineering & Resilience Testing

Design and execute chaos engineering experiments using Gremlin or Harness Chaos Engineering.

Simulate infrastructure, network, application, and dependency failures.

Validate system behavior during failure scenarios.

Establish reliability score baselines and improvement goals.

Measure resilience against real-world production conditions.

Document findings and implement corrective improvements.

 

Automation & Self-Healing

Identify repetitive operational tasks suitable for automation.

Develop self-healing workflows for common infrastructure and application failures.

Automate alert remediation, scaling, recovery, and operational activities.

Reduce manual intervention and operational toil.

Improve platform efficiency through engineering-driven automation.

Required Qualifications

8–10 years of experience in Site Reliability Engineering.

Experience with CSI for incident and RCA tracking.

Experience with Nobl9 for SLO management.

Experience with AppDynamics for application performance monitoring.

Experience with OpenTelemetry and Grafana Cloud for telemetry and observability.

Experience with Gremlin or Harness Chaos Engineering.

Experience with Akamai CDN.

Knowledge of AWS Well-Architected Framework.

Experience with Kubernetes reliability, observability, incident management, automation, and resilience engineering.

Regards,

Malya P

Lead Recruiter |Skanda Solutions

linkedin.com/in/malya-p-215dbs