1

Site Reliability Engineer Contract Jobs in Florida

Staff Site Reliability Engineer

Boca Raton, FL ยท On-site

$54 - $72/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards, and owns some of the most complex and consequential work on the platform. As a Staff Site ...

SRE/DevOps - Tampa, FL (Onsite)

Tampa, FL ยท On-site

$53.75 - $71.50/hr

Job Title: SRE/DevOps Location: Tampa, FL (Onsite) Duration: 1 Year Contract (Possible Extension) Experience: 8-12+ years of hands-on experience in Performance testing Minimum 4 Years of Banking ...

Senior Site Reliability Engineer II

Gainesville, FL ยท On-site +1

$125K - $209K/yr

We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role you will be ...

Site Reliability Engineering Lead

Boca Raton, FL ยท On-site

$54 - $72/hr

Lead Ops / SRE teams to build / maintain "Infrastructures as Code", software services (PaaS and SaaS), security policies and continuous integration / deployment processes * Remove technical debt ...

Senior Site Reliability Engineer II

Gainesville, FL ยท On-site +1

$125K - $209K/yr

We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role you will be ...

SRE Engineer

Jacksonville, FL ยท On-site

$52.75 - $70.25/hr

Jacksonville, FL, New York, NY (onsite) Job Type: full time SRE Engineer * Good knowledge on GCP * Hands on in defining and creation of CUJ, SLO, SLI, Error Budgeting based on NFR. * Strong Knowledge ...

Senior Site Reliability Engineer II

Boca Raton, FL ยท On-site +1

$125K - $209K/yr

We are hiring a hands-on Senior Site Reliability Engineer (SRE) to actively build, operate, and improve the reliability of our production systems. This is not a purely advisory role you will be ...

Site Reliability Engineer (AWS)

Jacksonville, FL ยท Hybrid

$52.75 - $70.25/hr

This position is under our CTO org to support SRE functions for innovation and growth for the Banking Solutions, Payments and Capital Markets business. This role will report under our Wealth and ...

Showing results 21-40

Site Reliability Engineer Contract information

What does a site reliability engineer contract do?

A Site Reliability Engineer (SRE) contractor is responsible for ensuring the reliability, scalability, and performance of software systems, typically on a temporary or project basis. They collaborate with development and operations teams to automate processes, monitor systems, and quickly resolve incidents. SRE contractors often design and implement tools that improve system uptime and efficiency, and may create documentation or best practices for site reliability. Their work helps organizations maintain high service availability while adapting to changing infrastructure needs.

What is the difference between Site Reliability Engineer Contract vs Site Reliability Engineer?

AspectSite Reliability Engineer ContractSite Reliability Engineer
CredentialsTypically requires SRE or related certifications, experience with cloud platformsSame as contract, often with more emphasis on full-time experience
Work EnvironmentProject-based, temporary, often remote or on-siteFull-time, ongoing, may be remote or on-site
Employer UsageUsed by companies for specific projects or to fill short-term needsUsed as a core role within organizations for continuous reliability management
Search & Comparison IntentCommonly compared for contract vs full-time roles in SREOften compared to contract roles for career planning

In summary, a Site Reliability Engineer Contract is a temporary, project-based role focusing on specific reliability tasks, while a full-time Site Reliability Engineer is a permanent position with ongoing responsibilities. Both roles require similar skills and certifications, but differ mainly in employment type and work setup.

What are the key skills and qualifications needed to thrive as a site reliability engineer contract?

Thriving as a Site Reliability Engineer Contractor requires strong expertise in systems administration, cloud platforms, automation, and programming (typically in Python, Go, or Bash), usually supported by a degree in computer science or relevant experience. Familiarity with DevOps tools like Kubernetes, Docker, Terraform, CI/CD pipelines, and monitoring solutions such as Prometheus or Grafana is essential, with certifications like AWS Certified DevOps Engineer or Google Professional SRE adding value. Exceptional problem-solving, collaboration, and communication skills help contractors quickly adapt to new environments and efficiently resolve incidents. These skills and qualities ensure reliable system performance, rapid response to outages, and seamless integration with client teams.

What are some common challenges faced by site reliability engineer contracts?

Site Reliability Engineers (SREs) on contract often face the challenge of quickly adapting to new systems, tooling, and organizational cultures. Since contracts are typically for a limited duration, contractors need to rapidly build rapport with internal teams and familiarize themselves with existing infrastructure to make impactful contributions. Additionally, they may need to balance multiple priorities, such as incident response, automation, and documentation, while ensuring that their work aligns with both short-term project goals and long-term site reliability. Effective communication and proactive collaboration with development and operations teams are essential for overcoming these challenges and delivering value within the contract period.
What are the most commonly searched types of Site Reliability Engineer jobs in Florida? The most popular types of Site Reliability Engineer jobs in Florida are:
What job categories do people searching Site Reliability Engineer Contract jobs in Florida look for? The top searched job categories for Site Reliability Engineer Contract jobs in Florida are:
What cities in Florida are hiring for Site Reliability Engineer Contract jobs? Cities in Florida with the most Site Reliability Engineer Contract job openings:
Infographic showing various Site Reliability Engineer Contract job openings in Florida as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 13% Part Time, 1% Temporary, and 3% Contract. Highlights an 94% Physical, 3% Hybrid, and 3% Remote job distribution.

Senior Site Reliability Engineer (SRE) - Observability & Resilience | Hybrid | Full-Time | Glendale,

SKANDA SOLUTIONS LLC

Florida City, FL โ€ข On-site

$52.50 - $69.75/hr

Other

This job post hasย expired 1 day ago.ย Applications are no longer accepted.


Job description

If you''re interested, please share your updated resume along with:

Current Location

Work Authorization

Expected Yearly full time Salary ?

JOB TITLE : Senior SRE

SKILL CATEGORY : Cloud: AWS

REQUIRED SKILLS : Site Reliability Engineering (SRE) & Kubernetes Operations

WORK LOCATION : (Glendale, Orlando, Seattle)- 3 Days Hybrid

Overview / Summary

We are seeking a Site Reliability Engineer (SRE) with 8โ€“10 years of experience to drive reliability, observability, and resilience improvements across critical systems. This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, continuous automation, and reliability engineering for Tier-1 business-critical applications.

Key Responsibilities

Drive automation initiatives to improve system performance and operational efficiency.

Improve application reliability and availability by proactively identifying and mitigating risks.

Analyze production incidents and root cause analyses (RCAs) to eliminate recurring issues and reduce outages.

Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets using Nobl9.

Conduct reliability assessments across applications, infrastructure, Kubernetes, databases, networks, caching platforms, and cloud environments.

Drive observability improvements using OpenTelemetry, Grafana Cloud, AppDynamics, Splunk, and monitoring best practices.

Perform performance and scalability reviews to support current and future demand.

Lead chaos engineering exercises using Gremlin or Harness Chaos Engineering.

Review cloud architectures against AWS Well-Architected Framework standards and drive remediation of reliability gaps.

Automate operational tasks and implement self-healing solutions.

Identify and eliminate single points of failure (SPOFs) and strengthen disaster recovery and failover capabilities.

Collaborate with Development, Infrastructure, Performance Engineering, and Operations teams to improve system resilience.

Establish reliability governance, dashboards, runbooks, and continuous improvement processes.

Reliability Assessment & Engineering

Conduct application reliability assessments using established reliability frameworks.

Review historical incidents, Sev-1/Sev-2 RCAs, and recurring failure patterns.

Identify reliability debt and drive remediation initiatives.

Evaluate application readiness for SRE engagement.

Perform end-to-end reliability reviews across application, infrastructure, network, and platform layers.

Define reliability roadmaps and track improvement initiatives.

Incident Management & RCA

Analyze incident trends using CSI or equivalent incident management platforms.

Participate in Major Incident Management and Problem Management processes.

Drive RCA reviews and corrective actions.

Track reliability improvement initiatives resulting from postmortems.

Reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

 

Service Level Management

Define and implement SLIs.

Establish SLOs and Error Budgets using Nobl9.

Partner with Product and Engineering teams to define business-focused reliability targets.

Build SLO dashboards and reliability scorecards.

Monitor error budget consumption and enforce governance policies.

Conduct reliability reviews based on SLO compliance.

Cloud & Platform Reliability

Review cloud architectures against AWS Well-Architected Framework principles.

Conduct reliability, performance, cost optimization, security, and operational excellence assessments.

Identify High Risk Issues (HRIs) and drive remediation.

Validate high availability, disaster recovery, backup, and failover capabilities.

Ensure multi-AZ and multi-region deployment strategies are implemented where required.

Kubernetes & Infrastructure Reliability

Review Kubernetes cluster health and workload configurations.

Validate resource requests, limits, autoscaling, and resiliency patterns.

Assess readiness, liveness, and startup probes.

Review service mesh configurations, network policies, and traffic routing.

Validate database high availability, caching strategies, and scaling configurations.

Identify and eliminate single points of failure.

Observability & Monitoring

Design and improve enterprise observability strategies.

Implement OpenTelemetry-based telemetry collection.

Manage metrics, events, logs, and traces (MELT).

Integrate telemetry into Grafana Cloud, Splunk Observability, or equivalent platforms.

Utilize AI-driven observability capabilities for anomaly detection and root cause analysis.

Improve alert quality, reduce alert fatigue, and increase actionable monitoring coverage.

Ensure every alert has an owner, runbook, and customer impact justification.

Application Performance Engineering

Conduct dependency mapping and architecture reviews.

Analyze latency, throughput, and scalability bottlenecks.

Review timeout, retry, circuit breaker, and resilience patterns.

Collaborate with Performance Engineering teams on load and stress testing.

Validate system capacity against current and future traffic demands.

Review Akamai CDN configurations, traffic routing, caching, and failover strategies.

Ensure applications can sustain significant traffic spikes and peak loads.

Chaos Engineering & Resilience Testing

Design and execute chaos engineering experiments using Gremlin or Harness Chaos Engineering.

Simulate infrastructure, network, application, and dependency failures.

Validate system behavior during failure scenarios.

Establish reliability score baselines and improvement goals.

Measure resilience against real-world production conditions.

Document findings and implement corrective improvements.

 

Automation & Self-Healing

Identify repetitive operational tasks suitable for automation.

Develop self-healing workflows for common infrastructure and application failures.

Automate alert remediation, scaling, recovery, and operational activities.

Reduce manual intervention and operational toil.

Improve platform efficiency through engineering-driven automation.

Required Qualifications

8โ€“10 years of experience in Site Reliability Engineering.

Experience with CSI for incident and RCA tracking.

Experience with Nobl9 for SLO management.

Experience with AppDynamics for application performance monitoring.

Experience with OpenTelemetry and Grafana Cloud for telemetry and observability.

Experience with Gremlin or Harness Chaos Engineering.

Experience with Akamai CDN.

Knowledge of AWS Well-Architected Framework.

Experience with Kubernetes reliability, observability, incident management, automation, and resilience engineering.

Regards,

Malya P

Lead Recruiter |Skanda Solutions

linkedin.com/in/malya-p-215dbs