1

Reliability Engineer Jobs in Miami, FL (NOW HIRING)

Site Reliability Engineer I

Sunrise, FL · On-site

$54.25 - $72.25/hr

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best ...

New

SRE AI Ops Engineer

Miami, FL · On-site

$54.50 - $72.50/hr

We are seeking an SRE AIOps Engineer to design, build, and support AI-driven automation solutions while providing advanced application and SRE operational support. This role combines AI engineering ...

New

Site Reliability Engineer I

Sunrise, FL · On-site

$78K - $124K/yr

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best ...

New

As a Senior Software Engineer in SRE at eMed, you will play a key role in ensuring our platform is highly available, secure, and performant. You'll lead reliability engineering efforts across ...

You'll lead reliability engineering efforts across production systems, drive operational excellence, and collaborate closely with application and infrastructure teams to design resilient services.

ENGINEER AUTOMATION RELIABILITY

Miami, FL · On-site

$87K - $110K/yr

The Reliability & Projects Engineer will be responsible for the identification, development, and implementation of innovative solutions to electrical/instrumentation/automation/networking reliability ...

ENGINEER AUTOMATION RELIABILITY

Miami, FL · On-site

$87K - $110K/yr

The Reliability & Projects Engineer will be responsible for the identification, development, and implementation of innovative solutions to electrical/instrumentation/automation/networking reliability ...

Software Engineer, Site Reliability

Miami, FL

$54.50 - $72.50/hr

Role As a Software Engineer working on Site Reliability at OpenEvidence, you will build and harden the mission-critical infrastructure powering our medical AI platform used by healthcare providers ...

Software Engineer, Site Reliability

Miami, FL · On-site

$54.50 - $72.50/hr

Role As a Software Engineer working on Site Reliability at OpenEvidence, you will build and harden the mission-critical infrastructure powering our medical AI platform used by healthcare providers ...

That combination -- early-stage velocity, regulatory weight, and high reliability expectations, requires an engineer whose primary mandate is reliability, scale, and operational excellence. This role ...

That combination -- early‑stage velocity, regulatory weight, and high reliability expectations, requires an engineer whose primary mandate is reliability, scale, and operational excellence. This ...

That combination -- early-stage velocity, regulatory weight, and high reliability expectations, requires an engineer whose primary mandate is reliability, scale, and operational excellence. This role ...

That combination - early-stage velocity, regulatory weight, and high reliability expectations, requires an engineer whose primary mandate is reliability, scale, and operational excellence. This role ...

That combination - early-stage velocity, regulatory weight, and high reliability expectations, requires an engineer whose primary mandate is reliability, scale, and operational excellence. This role ...

Showing results 21-40

Reliability Engineer information

See Miami, FL salary details

$58.3K

$112.8K

$134.9K

How much do reliability engineer jobs pay per year?

As of Sep 3, 2026, the average yearly pay for reliability engineer in Miami, FL is $112,834.00, according to ZipRecruiter salary data. Most workers in this role earn between $98,000.00 and $123,400.00 per year, depending on experience, location, and employer.

What is a reliability engineer?

Reliability Engineers are professionals responsible for ensuring that systems, equipment, or processes function consistently and efficiently over time. They analyze data, identify potential points of failure, and develop maintenance strategies to improve system reliability and minimize downtime. Their work spans various industries, including manufacturing, energy, and technology, and often involves collaborating with design, operations, and maintenance teams. By implementing reliability-centered maintenance and predictive analysis, they help organizations save costs and increase safety.

What does a reliability engineer do?

As a reliability engineer, your duties are to test and evaluate the manufacturing of products and components and ensure that the procedures are efficient and do not lead to abnormally high maintenance or operational costs. Your other responsibilities are to find solutions to product reliability risks. You may manage risk in a supply chain, develop loss prevention strategies, and track the entire lifecycle of product development, from building prototypes to moving a product into full-scale production. You analyze information from department heads and recommend strategies to reduce risk and ensure that the product works reliably.

What are the key skills and qualifications needed to thrive as a reliability engineer, and why are they important?

To thrive as a Reliability Engineer, you need a solid background in engineering principles, failure analysis, and reliability modeling, typically with a degree in engineering or a related field. Familiarity with tools such as FMEA, Root Cause Analysis (RCA), reliability-centered maintenance (RCM) software, and certifications like Certified Reliability Engineer (CRE) are highly valued. Strong problem-solving abilities, attention to detail, and effective communication are crucial soft skills in this role. These skills ensure systems are dependable, downtime is minimized, and organizational performance and safety are optimized.

What are some typical challenges reliability engineers face when implementing preventive maintenance strategies?

Reliability Engineers often encounter challenges such as balancing preventive maintenance schedules with production demands, ensuring buy-in from operations teams, and accurately predicting equipment failures. They must analyze large sets of historical data to identify trends and root causes, which can be complex in facilities with diverse machinery. Collaboration with maintenance, operations, and engineering teams is essential to develop effective strategies that minimize downtime while optimizing resources.

What is the difference between Reliability Engineer vs Maintenance Engineer?

AspectReliability EngineerMaintenance Engineer
CredentialsTypically requires engineering degree, certifications in reliability or asset managementOften requires engineering or technical diploma, certifications in maintenance or equipment repair
Work EnvironmentFocuses on analysis, design, and improvement of systems for reliabilityHands-on maintenance, repair, and troubleshooting of equipment
Industry UsageCommon in manufacturing, energy, aerospace, and industrial sectorsPrevalent in manufacturing, facilities, and industrial plants

Reliability Engineers focus on designing and improving systems to prevent failures, using data analysis and modeling. Maintenance Engineers perform hands-on repairs and upkeep of equipment to ensure operational continuity. While both roles aim to optimize equipment performance, Reliability Engineers work proactively on system reliability, whereas Maintenance Engineers handle reactive and scheduled maintenance tasks.

Are reliability engineers in demand?

Reliability engineers are in high demand across industries such as manufacturing, energy, and aerospace due to their role in improving system performance and reducing downtime. Employers seek professionals with skills in data analysis, failure modes, and maintenance strategies, often requiring certifications like Certified Reliability Engineer (CRE). The job outlook is positive, with steady growth expected as companies prioritize operational efficiency and risk management.

How much do reliability engineers get paid?

Reliability engineers typically earn a median annual salary ranging from $70,000 to $110,000, depending on experience, location, and industry. Senior or specialized reliability engineers with certifications and advanced skills can earn higher salaries, often exceeding $120,000 annually.

What are the most commonly searched types of Reliability Engineer jobs in Miami, FL?

The most popular types of Reliability Engineer jobs in Miami, FL are:

What job categories do people searching Reliability Engineer jobs in Miami, FL look for?

The top searched job categories for Reliability Engineer jobs in Miami, FL are:

What cities near Miami, FL are hiring for Reliability Engineer jobs?

Cities near Miami, FL with the most Reliability Engineer job openings:

Infographic showing various Reliability Engineer job openings in Miami, FL as of August 2026, with employment types broken down into 100% Full Time. Highlights an 74% In-person, and 26% Remote job distribution, with an average salary of $112,834 per year, or $54.2 per hour.

Site Reliability Engineer I

American Express

Sunrise, FL • On-site

$54.25 - $72.25/hr

Full-time

Posted 2 days ago

New


American Express rating

8.6

Company rating: 8.6 out of 10

Based on 37 frontline employees who took The Breakroom Quiz

25th of 152 rated financial services


Job description

Site Reliability Engineer I enhances system resilience and performance, implements automation tools, and contributes to the architectural design and disaster recovery strategies, promoting best practices for continuous improvement and reliability.

Education Qualifications:

  • Minimum of 5+ years of relevant experience in application development, maintenance, and production support, along with hands-on exposure to Java and distributed systems in enterprise environments.
  • Bachelor's degree in computer science, Information Technology, Engineering, or equivalent practical experience; advanced degree is a plus
  • Strong knowledge of operating systems and application runtimes such as Java and .NET
  • Knowledge of distributed systems and servicebased architectures from an operations and reliability perspective
  • Strong knowledge of modern observability stacks and platforms, including Splunk, Elasticsearch, Prometheus, and Grafana
  • Knowledge of observability practices including logging, monitoring, tracing, and performance analysis
  • Knowledge of RDBMS and NoSQL databases including MySQL, PostgreSQL, Couchbase, HBase, and Cassandra
  • Knowledge of scripting and automation using languages such as PowerShell and Python
  • Knowledge of AI, analytics, or AIOps platforms from an operational perspective is a plus

Work Experience:

  • Experience in Incident, Problem, and Change Management using ServiceNow or similar ITSM tools
  • Experience supporting production systems in largescale enterprise environments with a focus on reliability and availability
  • Experience in system administration, infrastructure operations, and network troubleshooting
  • Experience with CI/CD pipeline implementation and support using tools such as Jenkins, GitHub Actions, XL Release (XLR), or similar
  • Experience managing and troubleshooting technology infrastructure and services, including servers, networks, and cloud platforms
  • Knowledge of cloudbased Site Reliability Engineering (SRE) practices with handson experience on public cloud platforms such as AWS, Azure, or Google Cloud Platform
  • Knowledge of containerization and orchestration technologies such as Docker and Kubernetes, and microservicesbased architectures
  • Experience using enterprise monitoring and alerting platforms such as ELF
  • Exposure to AIassisted monitoring, automation, or AIOps tools is a plus
    Proficiency in connecting to and administering servers via SSH (Secure Shell)
  • Knowledge of core networking concepts including ports, protocols, firewalls, and secure remote access

Licenses & Certifications

  • Certification in at least one programming language or runtime such as Java, .NET, or Python
  • Certification in containerization and orchestration technologies (Docker, Kubernetes, OpenShift) is a plus
  • Public cloud certification in AWS or GCP is a plus
  • Certification or training related to AI platforms, analytics platforms, or AIOps is a plus

Employment eligibility to work with American Express in the United States is required as the company will not pursue visa sponsorship for these positions. 

  • Monitor application and infrastructure health using enterprise monitoring and observability tools, including ELF, to ensure availability, performance, and reliability of enterprise platforms
  • Configure, tune, and maintain alerting mechanisms in ELF, aligned to service health indicators and SLOs, to enable timely incident detection and reduce noise and false positives
  • Develop and maintain dashboards providing visibility into system performance, availability, reliability trends, and key operational metrics
  • Analyze metrics, logs, and distributed traces across application and infrastructure layers to proactively identify issues and support effective root cause analysis (RCA)
  • Own and execute blameless RCAs for production incidents, identify corrective and preventive actions, and track them to closure
  • Implement minor code fixes, configuration updates, and reliability enhancements as part of incident remediation and preventive measures
  • Collaborate with application development and platform teams to review defects, propose fixes, and improve overall service reliability
  • Participate in Agile sprint planning ceremonies, backlog grooming, estimation, and delivery of SREowned work items
  • Drive reliability improvements through sprintbased commitments, including automation, operational fixes, and platform enhancements
  • Participate in Disaster Recovery (DR) planning, testing, and execution to ensure resilience of businesscritical services
  • Perform regular system patching and maintenance activities in line with organizational security, compliance, and audit requirements
  • Support ITILbased Incident, Problem, and Change Management processes, including planning, documentation, approvals, execution, and postimplementation validation
  • Monitor network performance and troubleshoot connectivity, latency, and accessrelated issues impacting platform traffic
  • Participate in certificate lifecycle management, including provisioning, renewal, validation, and troubleshooting of SSL/TLS certificates
  • Maintain and manage service accounts (Service IDs), including access provisioning, credential rotation, and compliance with security policies
  • Drive automation and operational toil reduction using scripting, CI/CD pipelines, and platform tooling to improve reliability and scalability
  • Maintain accurate documentation of system configurations, runbooks, SOPs, platform operational guidelines, and troubleshooting procedures, and generate reports on system performance, incidents, and resolutions
  • Participate and lead the Development change review and change validation processes
  • Collaborates with senior engineers to contribute to the architectural design of systems, ensuring that reliability, scalability, and performance considerations are integrated into design discussions with direct guidance from senior colleagues
  • Uses AI-assisted coding and documentation tools to support development of automation scripts, runbooks, and infrastructure as code with guidance from senior engineers

What American Express employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom