1

Site Reliability Engineer Contract Jobs in Florida

Site Reliability Engineer I

Sunrise, FL · On-site

$54.25 - $72.25/hr

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best ...

New

Site Reliability Engineer II

Orlando, FL

$53.25 - $70.75/hr

Site Reliability Engineer II The SRE II sits at the intersection of software engineering and platform operations. You will own the reliability, scalability, and operational hygiene of Kastle's core ...

SRE Engineer

Jacksonville, FL · On-site

$52.75 - $70.25/hr

Jacksonville, FL, New York, NY (onsite) Job Type: full time SRE Engineer * Good knowledge on GCP * Hands on in defining and creation of CUJ, SLO, SLI, Error Budgeting based on NFR. * Strong Knowledge ...

Lead Ops / SRE teams to build / maintain "Infrastructures as Code", software services (PaaS and SaaS), security policies and continuous integration / deployment processes * Remove technical debt ...

Leidos is seeking a Site Reliability Engineer (SRE) Data Engineer supporting the largest IT services program for the Navy. Under the Service Management, Integration, and Transport (SMIT) program, the ...

Site Reliability Data Engineer

Jacksonville, FL · On-site

$52.75 - $70.25/hr

Leidos is seeking a Site Reliability Engineer (SRE) Data Engineer supporting the largest IT services program for the Navy. Under the Service Management, Integration, and Transport (SMIT) program, the ...

Senior Site Reliability Engineer Location: Orlando HQ - Remote, FL Job Id: 615 # of Openings: 1 Senior Site Reliability Engineer Location: Remote Compensation: $160,000 - 180,000 per year, depending ...

Site Reliability Engineer - Onsite

Tampa, FL · On-site

$53.75 - $71.50/hr

Site Reliability Engineer Work Location: Jersey City, NJ / Tampa, FL Duration: Long Term * Design, implement, and maintain scalable and resilient Apache Flink deployments on Kubernetes. * Develop ...

Site Reliability Engineer I

Sunrise, FL · On-site

$78K - $124K/yr

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best ...

Showing results 41-60

Site Reliability Engineer Contract information

What does a site reliability engineer contract do?

A Site Reliability Engineer (SRE) contractor is responsible for ensuring the reliability, scalability, and performance of software systems, typically on a temporary or project basis. They collaborate with development and operations teams to automate processes, monitor systems, and quickly resolve incidents. SRE contractors often design and implement tools that improve system uptime and efficiency, and may create documentation or best practices for site reliability. Their work helps organizations maintain high service availability while adapting to changing infrastructure needs.

What are the key skills and qualifications needed to thrive as a site reliability engineer contract?

Thriving as a Site Reliability Engineer Contractor requires strong expertise in systems administration, cloud platforms, automation, and programming (typically in Python, Go, or Bash), usually supported by a degree in computer science or relevant experience. Familiarity with DevOps tools like Kubernetes, Docker, Terraform, CI/CD pipelines, and monitoring solutions such as Prometheus or Grafana is essential, with certifications like AWS Certified DevOps Engineer or Google Professional SRE adding value. Exceptional problem-solving, collaboration, and communication skills help contractors quickly adapt to new environments and efficiently resolve incidents. These skills and qualities ensure reliable system performance, rapid response to outages, and seamless integration with client teams.

What are some common challenges faced by site reliability engineer contracts?

Site Reliability Engineers (SREs) on contract often face the challenge of quickly adapting to new systems, tooling, and organizational cultures. Since contracts are typically for a limited duration, contractors need to rapidly build rapport with internal teams and familiarize themselves with existing infrastructure to make impactful contributions. Additionally, they may need to balance multiple priorities, such as incident response, automation, and documentation, while ensuring that their work aligns with both short-term project goals and long-term site reliability. Effective communication and proactive collaboration with development and operations teams are essential for overcoming these challenges and delivering value within the contract period.

What is the difference between Site Reliability Engineer Contract vs Site Reliability Engineer?

AspectSite Reliability Engineer ContractSite Reliability Engineer
CredentialsTypically requires SRE or related certifications, experience with cloud platformsSame as contract, often with more emphasis on full-time experience
Work EnvironmentProject-based, temporary, often remote or on-siteFull-time, ongoing, may be remote or on-site
Employer UsageUsed by companies for specific projects or to fill short-term needsUsed as a core role within organizations for continuous reliability management
Search & Comparison IntentCommonly compared for contract vs full-time roles in SREOften compared to contract roles for career planning

In summary, a Site Reliability Engineer Contract is a temporary, project-based role focusing on specific reliability tasks, while a full-time Site Reliability Engineer is a permanent position with ongoing responsibilities. Both roles require similar skills and certifications, but differ mainly in employment type and work setup.

What are the most commonly searched types of Site Reliability Engineer jobs in Florida?

The most popular types of Site Reliability Engineer jobs in Florida are:

What are popular job titles related to Site Reliability Engineer Contract jobs in Florida?

For Site Reliability Engineer Contract jobs in Florida, the most frequently searched job titles are:

What job categories do people searching Site Reliability Engineer Contract jobs in Florida look for?

The top searched job categories for Site Reliability Engineer Contract jobs in Florida are:

What cities in Florida are hiring for Site Reliability Engineer Contract jobs?

Cities in Florida with the most Site Reliability Engineer Contract job openings:

Infographic showing various Site Reliability Engineer Contract job openings in Florida as of August 2026, with employment types broken down into 62% Full Time, and 38% Contract. Highlights an 77% In-person, and 23% Remote job distribution.

Site Reliability Engineer I

American Express

Sunrise, FL • On-site

$54.25 - $72.25/hr

Full-time

Posted 3 days ago

New


American Express rating

8.6

Company rating: 8.6 out of 10

Based on 37 frontline employees who took The Breakroom Quiz

25th of 152 rated financial services


Job description

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.

Education Qualifications:

  • Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.
  • Bachelor's degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus
  • Strong knowledge of operating systems and application runtimes such as Java and .NET
  • Knowledge of distributed systems and servicebased architectures from an operations and reliability perspective
  • Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana
  • Knowledge of observability practices including logging, monitoring, tracing, and performance analysis
  • Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra
  • Knowledge of scripting and automation using languages such as PowerShell and Python
  • Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus

Work Experience:

  • Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools
  • Experience supporting production systems in largescale enterprise environments with a focus on reliability and availability
  • Experience in system administration, infrastructure operations, and network troubleshooting
  • Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar
  • Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms
  • Knowledge of cloudbased Site Reliability Engineering (SRE) practices with handson experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform
  • Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservicesbased architectures
  • Experience using enterprise monitoring and alerting platforms such as ELF
  • Exposure to AIassisted monitoring, automation, or AIOps tools is a plus
    Proficiency in connecting to and administering servers via SSH (Secure Shell)
  • Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access

Licenses & Certifications

  • Certification in at least one programming language or runtime such as Java, .NET, or Python
  • Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus
  • Public cloud certification in AWS or GCP is a plus
  • Certification or training related to AI platforms, analytics platforms, or AIOps is a plus

Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions. 

  • Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms
  • Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives
  • Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics
  • Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)
  • Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure
  • Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures
  • Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability
  • Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SREowned work items
  • Drive reliability improvements through sprintbased commitments, including automation, operational fixes, and platform enhancements
  • Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of businesscritical services
  • Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements
  • Support ITILbased Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and postimplementation validation
  • Monitor network performance and troubleshoot connectivity, latency, and accessrelated issues impacting platform traffic
  • Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates
  • Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies
  • Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability
  • Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions
  • Participate and lead the Development change review and change validation processes
  • Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
  • Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers

What American Express employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom