1

Network Reliability Engineer Jobs in Miami, FL (NOW HIRING)

ENGINEER AUTOMATION RELIABILITY

Miami, FL · On-site

$87K - $110K/yr

Bachelor's Degree in Electrical, or Electronic, or Computer, or Networking Engineering with a minimum of 12 years of experience in areas of engineering, maintenance, reliability, automation ...

ENGINEER AUTOMATION RELIABILITY

Miami, FL

$87K - $110K/yr

Bachelor's Degree in Electrical, or Electronic, or Computer, or Networking Engineering with a minimum of 12 years of experience in areas of engineering, maintenance, reliability, automation ...

Site Reliability Engineer

Miami, FL

$54.50 - $72.50/hr

Create and maintain reusable platform abstractions across AWS and Azure that standardize security, reliability, networking, and observability. * Reduce developer cognitive load by abstracting ...

Senior Software Engineer (SRE)

Miami, FL · On-site

$54.50 - $72.25/hr

The role involves leading reliability engineering efforts, collaborating with teams to design ... network controls, secrets management, and compliance readiness • Participate in and lead in ...

You'll lead reliability engineering efforts across production systems, drive operational excellence ... Contribute to security hardening efforts, including network controls, secrets management, and ...

Familiarity with Ceph storage, OVN/OVS networking, and observability tools such as Prometheus, Grafana, or Zabbix. * Experience building SRE practices for platform services, including SLIs/SLOs ...

NETWORK ENGINEER

Opa Locka, FL · On-site

$60K - $78K/yr

The Network Engineer ensures high availability, reliability, security, and performance of networked systems, while working closely with field teams, civil/electrical engineers, vendor partners, DOTs ...

DevOps Engineer

Miami, FL

$51 - $70/hr

Proven experience in DevOps/SRE position * Deep understanding of Linux basics and containers ... Networking knowledge, according to CCNA level * Backup and restore strategies, disaster recovery ...

DevOps Engineer

Miami, FL · On-site

$51 - $70/hr

Proven experience in DevOps/SRE position * Deep understanding of Linux basics and containers ... Networking knowledge, according to CCNA level * Backup and restore strategies, disaster recovery ...

DevOps Engineer

Miami, FL · On-site

$51 - $70/hr

Proven experience in DevOps/SRE position * Deep understanding of Linux basics and containers ... Networking knowledge, according to CCNA level * Backup and restore strategies, disaster recovery ...

... reliability. * Assist in the implementation of cloud connectivity solutions, optimizing for ... in both network engineering and telecommunications, gaining hands-on experience across ...

... reliability. Assist in the implementation of cloud connectivity solutions, optimizing for ... in both network engineering and telecommunications, gaining hands-on experience across ...

... reliability. * Assist in the implementation of cloud connectivity solutions, optimizing for ... in both network engineering and telecommunications, gaining hands-on experience across ...

Senior Network Engineer Doral, Florida Employment Type: Full-Time Employee preferred; 1099 may be ... Provide senior-level technical leadership and recommendations to improve reliability, security, and ...

next page

Showing results 1-20

Network Reliability Engineer information

See Miami, FL salary details

$58.3K

$112.8K

$134.9K

How much do network reliability engineer jobs pay per year?

As of Jul 27, 2026, the average yearly pay for network reliability engineer in Miami, FL is $112,834.00, according to ZipRecruiter salary data. Most workers in this role earn between $98,000.00 and $123,400.00 per year, depending on experience, location, and employer.

What is a Network Reliability Engineer?

A Network Reliability Engineer (NRE) is an IT professional responsible for ensuring the reliability, performance, and scalability of network systems. They combine skills in networking, software engineering, and automation to proactively detect and resolve potential network issues before they affect users. NREs often design and implement monitoring tools, automate network management tasks, and work to improve the overall stability of network infrastructure. Their goal is to minimize downtime and ensure seamless connectivity across an organization’s network.

What is the difference between Network Reliability Engineer vs Network Operations Center (NOC) Technician?

AspectNetwork Reliability EngineerNetwork Operations Center (NOC) Technician
CertificationsCCNA, CCNP, Network+CCNA, Network+
Work EnvironmentDesign, analyze, and improve network infrastructureMonitor, troubleshoot, and maintain networks in real-time
Employer & Industry UsageTelecom, large enterprises, cloud providersISPs, data centers, enterprise networks
Common Search & ComparisonFocus on network reliability and designFocus on network monitoring and incident response

The main difference is that Network Reliability Engineers focus on designing and improving network systems to ensure long-term reliability, while NOC Technicians monitor and troubleshoot networks in real-time to resolve issues quickly. Both roles require relevant certifications and are essential in maintaining network performance, but they serve different functions within network management.

What are the key skills and qualifications needed to thrive as a Network Reliability Engineer, and why are they important?

To thrive as a Network Reliability Engineer, you need a strong background in computer networking, network protocols, troubleshooting, and often a degree in computer science or a related field. Familiarity with tools like Wireshark, Nagios, Cisco IOS, and certifications such as CCNA or CCNP are commonly required. Analytical thinking, proactive problem-solving, and effective communication are standout soft skills in this role. These skills are crucial to maintaining reliable network operations, minimizing downtime, and ensuring seamless communication across organizational systems.

What are some common challenges faced by Network Reliability Engineers, and how are they typically addressed?

Network Reliability Engineers often encounter challenges such as diagnosing intermittent connectivity issues, managing network upgrades with minimal downtime, and maintaining high availability during peak traffic. These are typically addressed by leveraging robust monitoring tools, implementing automation for routine tasks, and collaborating closely with software engineers, network administrators, and incident response teams. Staying current with evolving network technologies and best practices is essential for effectively identifying and resolving problems before they impact users.
What are popular job titles related to Network Reliability Engineer jobs in Miami, FL? For Network Reliability Engineer jobs in Miami, FL, the most frequently searched job titles are:
What job categories do people searching Network Reliability Engineer jobs in Miami, FL look for? The top searched job categories for Network Reliability Engineer jobs in Miami, FL are:
What cities near Miami, FL are hiring for Network Reliability Engineer jobs? Cities near Miami, FL with the most Network Reliability Engineer job openings:
Infographic showing various Network Reliability Engineer job openings in Miami, FL as of June 2026, with employment types broken down into 80% Full Time, and 20% Contract. Highlights an 100% In-person job distribution, with an average salary of $112,834 per year, or $54.2 per hour.

Senior Site Reliability Engineer (SRE) - Observability & Resilience | Hybrid | Full-Time | Glendale,

SKANDA SOLUTIONS LLC

Florida City, FL • On-site

$52.50 - $69.75/hr

Other

Posted 18 days ago


Job description

If you''re interested, please share your updated resume along with:

Current Location

Work Authorization

Expected Yearly full time Salary ?

JOB TITLE : Senior SRE

SKILL CATEGORY : Cloud: AWS

REQUIRED SKILLS : Site Reliability Engineering (SRE) & Kubernetes Operations

WORK LOCATION : (Glendale, Orlando, Seattle)- 3 Days Hybrid

Overview / Summary

We are seeking a Site Reliability Engineer (SRE) with 8–10 years of experience to drive reliability, observability, and resilience improvements across critical systems. This is a high-impact, front-line operations role focused on real-time incident response, proactive prevention, continuous automation, and reliability engineering for Tier-1 business-critical applications.

Key Responsibilities

Drive automation initiatives to improve system performance and operational efficiency.

Improve application reliability and availability by proactively identifying and mitigating risks.

Analyze production incidents and root cause analyses (RCAs) to eliminate recurring issues and reduce outages.

Define and manage Service Level Objectives (SLOs), Service Level Indicators (SLIs), and Error Budgets using Nobl9.

Conduct reliability assessments across applications, infrastructure, Kubernetes, databases, networks, caching platforms, and cloud environments.

Drive observability improvements using OpenTelemetry, Grafana Cloud, AppDynamics, Splunk, and monitoring best practices.

Perform performance and scalability reviews to support current and future demand.

Lead chaos engineering exercises using Gremlin or Harness Chaos Engineering.

Review cloud architectures against AWS Well-Architected Framework standards and drive remediation of reliability gaps.

Automate operational tasks and implement self-healing solutions.

Identify and eliminate single points of failure (SPOFs) and strengthen disaster recovery and failover capabilities.

Collaborate with Development, Infrastructure, Performance Engineering, and Operations teams to improve system resilience.

Establish reliability governance, dashboards, runbooks, and continuous improvement processes.

Reliability Assessment & Engineering

Conduct application reliability assessments using established reliability frameworks.

Review historical incidents, Sev-1/Sev-2 RCAs, and recurring failure patterns.

Identify reliability debt and drive remediation initiatives.

Evaluate application readiness for SRE engagement.

Perform end-to-end reliability reviews across application, infrastructure, network, and platform layers.

Define reliability roadmaps and track improvement initiatives.

Incident Management & RCA

Analyze incident trends using CSI or equivalent incident management platforms.

Participate in Major Incident Management and Problem Management processes.

Drive RCA reviews and corrective actions.

Track reliability improvement initiatives resulting from postmortems.

Reduce Mean Time to Detect (MTTD) and Mean Time to Recover (MTTR).

 

Service Level Management

Define and implement SLIs.

Establish SLOs and Error Budgets using Nobl9.

Partner with Product and Engineering teams to define business-focused reliability targets.

Build SLO dashboards and reliability scorecards.

Monitor error budget consumption and enforce governance policies.

Conduct reliability reviews based on SLO compliance.

Cloud & Platform Reliability

Review cloud architectures against AWS Well-Architected Framework principles.

Conduct reliability, performance, cost optimization, security, and operational excellence assessments.

Identify High Risk Issues (HRIs) and drive remediation.

Validate high availability, disaster recovery, backup, and failover capabilities.

Ensure multi-AZ and multi-region deployment strategies are implemented where required.

Kubernetes & Infrastructure Reliability

Review Kubernetes cluster health and workload configurations.

Validate resource requests, limits, autoscaling, and resiliency patterns.

Assess readiness, liveness, and startup probes.

Review service mesh configurations, network policies, and traffic routing.

Validate database high availability, caching strategies, and scaling configurations.

Identify and eliminate single points of failure.

Observability & Monitoring

Design and improve enterprise observability strategies.

Implement OpenTelemetry-based telemetry collection.

Manage metrics, events, logs, and traces (MELT).

Integrate telemetry into Grafana Cloud, Splunk Observability, or equivalent platforms.

Utilize AI-driven observability capabilities for anomaly detection and root cause analysis.

Improve alert quality, reduce alert fatigue, and increase actionable monitoring coverage.

Ensure every alert has an owner, runbook, and customer impact justification.

Application Performance Engineering

Conduct dependency mapping and architecture reviews.

Analyze latency, throughput, and scalability bottlenecks.

Review timeout, retry, circuit breaker, and resilience patterns.

Collaborate with Performance Engineering teams on load and stress testing.

Validate system capacity against current and future traffic demands.

Review Akamai CDN configurations, traffic routing, caching, and failover strategies.

Ensure applications can sustain significant traffic spikes and peak loads.

Chaos Engineering & Resilience Testing

Design and execute chaos engineering experiments using Gremlin or Harness Chaos Engineering.

Simulate infrastructure, network, application, and dependency failures.

Validate system behavior during failure scenarios.

Establish reliability score baselines and improvement goals.

Measure resilience against real-world production conditions.

Document findings and implement corrective improvements.

 

Automation & Self-Healing

Identify repetitive operational tasks suitable for automation.

Develop self-healing workflows for common infrastructure and application failures.

Automate alert remediation, scaling, recovery, and operational activities.

Reduce manual intervention and operational toil.

Improve platform efficiency through engineering-driven automation.

Required Qualifications

8–10 years of experience in Site Reliability Engineering.

Experience with CSI for incident and RCA tracking.

Experience with Nobl9 for SLO management.

Experience with AppDynamics for application performance monitoring.

Experience with OpenTelemetry and Grafana Cloud for telemetry and observability.

Experience with Gremlin or Harness Chaos Engineering.

Experience with Akamai CDN.

Knowledge of AWS Well-Architected Framework.

Experience with Kubernetes reliability, observability, incident management, automation, and resilience engineering.

Regards,

Malya P

Lead Recruiter |Skanda Solutions

linkedin.com/in/malya-p-215dbs