1

Software Reliability Engineer Jobs (NOW HIRING)

Lead Software Engineer - Reliability

Miami, FL · On-site

$100 - $130/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

... weight, and high reliability expectations, requires an engineer whose primary mandate is ... Identify repetitive operational work and eliminate it with software -- automation, self‑healing ...

Sr. Software Engineer (Flight Reliability)

Hawthorne, CA · On-site +1

$160K - $225K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

SR. SOFTWARE ENGINEER (FLIGHT RELIABILITY) The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and ...

Reliability Engineer

Franklin, OH · On-site

$94K - $118K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

The Reliability Engineer is responsible for improving the reliability, availability, and ... Use data analysis tools and reliability software to model performance and predict failures.

Reliability Engineer

Franklin, OH

$94K - $118K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

The Reliability Engineer is responsible for improving the reliability, availability, and ... Use data analysis tools and reliability software to model performance and predict failures.

Software Engineer (Flight Reliability)

Hawthorne, CA · On-site +1

$145K - $175K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

SOFTWARE ENGINEER (FLIGHT RELIAIBLITY) The Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and flight as ...

$140 - $190/hr

Formal training or certification in software engineering concepts plus 5 years of applied experience * Proficiency in reliability, scalability, performance, security, toil reduction and site ...

New

Lead Software Engineer - Reliability

Miami, FL · On-site

$140 - $180/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

... weight, and high reliability expectations, requires an engineer whose primary mandate is ... Identify repetitive operational work and eliminate it with software -- automation, self-healing ...

.NET Software Engineer / SRE II

Irving, TX · On-site

$54.75 - $72.75/hr

  • Dental

  • Vision

  • Retirement

.NET Software Engineer / SRE II DETAILS Location : Location : Arlington, TX 76014 (hybrid onsite 2-days per week) Position Type : Direct-Hire Hourly / Salary : to $145K (based on experience) + 10% ...

Lead Software Engineer - Reliability

Miami, FL · On-site

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

... weight, and high reliability expectations, requires an engineer whose primary mandate is ... Identify repetitive operational work and eliminate it with software - automation, self-healing ...

Reliability Engineer

Irvine, CA · On-site

$110K - $138K/yr

Reliability Engineer Full Time 40 hours/Week Duration: 12 months and flexible to extend further ... Proficiency in relevant software/tools (reliability modelling software, statistical tools, MS Excel ...

Software Engineer III- SRE

Wilmington, DE · On-site

  • Medical

  • Retirement

Formal training or certification in software engineering concepts plus 5 years of applied experience * Proficiency in reliability, scalability, performance, security, toil reduction and site ...

Showing results 41-60

Software Reliability Engineer information

See salary details

$39

$67

$88

How much do software reliability engineer jobs pay per hour?

As of Aug 13, 2026, the average hourly pay for software reliability engineer in the United States is $67.07, according to ZipRecruiter salary data. Most workers in this role earn between $59.13 and $74.52 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a software reliability engineer, and why are they important?

To thrive as a Software Reliability Engineer, you need a strong background in software development, system architecture, and incident response, often supported by a degree in computer science or related field. Familiarity with monitoring tools (like Prometheus), cloud platforms (AWS, GCP), automation frameworks, and certifications such as AWS Certified DevOps Engineer are highly valuable. Excellent problem-solving, collaboration, and communication skills help you coordinate effectively during high-pressure situations and with cross-functional teams. These abilities are crucial for maintaining system uptime, quickly resolving outages, and ensuring the overall reliability of critical software services.

What is a software reliability engineer?

Software Reliability Engineers (SREs) are IT professionals who focus on ensuring that software systems are reliable, scalable, and maintain high availability. They work at the intersection of software development and IT operations, often automating processes, monitoring system performance, and responding to incidents. SREs use engineering principles to solve operational problems, aiming to reduce downtime and improve user experience. Their responsibilities can include building tools, managing infrastructure, and collaborating with development teams to implement best practices for reliability.

How does a software reliability engineer typically interact with development and operations teams to improve system stability?

Software Reliability Engineers (SREs) work closely with both development and operations teams to ensure that systems are reliable, scalable, and maintainable. They often participate in design reviews, provide input on architectural decisions, and help define service-level objectives. SREs also collaborate with developers to automate deployment processes and create monitoring solutions, and they partner with operations staff to manage incident response and root cause analysis. This collaborative environment enables them to proactively identify potential issues and drive cross-functional improvements.

What is the difference between Software Reliability Engineer vs Software Test Engineer?

AspectSoftware Reliability EngineerSoftware Test Engineer
Primary FocusEnsuring software reliability, stability, and performance over timeDesigning and executing tests to identify bugs and verify functionality
Skills & CertificationsKnowledge of reliability engineering, scripting, monitoring toolsTesting methodologies, automation tools, scripting
Work EnvironmentCollaborates with development and operations teams, often in DevOpsWorks primarily in QA/testing teams, often in dedicated testing phases
Industry UsageCommon in software companies focusing on product stabilityWidely used in software development and QA departments

The main difference is that Software Reliability Engineers focus on maintaining long-term software stability and performance, while Software Test Engineers concentrate on identifying bugs through testing. Both roles require technical skills and often collaborate, but their core objectives differ: reliability versus defect detection.

More about Software Reliability Engineer jobs
What cities are hiring for Software Reliability Engineer jobs? Cities with the most Software Reliability Engineer job openings:
Who are the top companies hiring for Software Reliability Engineer jobs? The top employers for Software Reliability Engineer jobs are:
Infographic showing various Software Reliability Engineer job openings in the United States as of August 2026, with employment types broken down into 25% Internship, and 75% Full Time. Highlights an 100% In-person job distribution, with an average salary of $139,500 per year, or $67.1 per hour.

Lead Software Engineer - Reliability

Socket.dev

Miami, FL • On-site

$100 - $130/hr

Other

Medical, Dental, Vision, Life, Retirement

Posted 7 days ago


Job description

Nu is one of the largest digital financial platforms in the world, with more than 127 million customers across Brazil, Mexico, and Colombia. Guided by our mission to fight complexity and empower people, we are redefining financial services in Latin America and this is still just the beginning of the purple future we're building.

Listed on the New York Stock Exchange (NYSE: NU), we combine proprietary technology, data intelligence, and an efficient operating model to deliver financial products that are simple, accessible, and human.

Our impact has been recognized by global rankings such as Time 100 Companies, Fast Company’s Most Innovative Companies, and Forbes World’s Best Bank. Visit our institutional page https://international.nubank.com.br/careers/

About the role

The U.S. Market team is launching a differentiated financial product in the largest and most demanding financial market in the world. We’re iterating quickly on real customer signals while building systems that will eventually serve customers at Nubank scale. That combination — early‑stage velocity, regulatory weight, and high reliability expectations, requires an engineer whose primary mandate is reliability, scale, and operational excellence.

This role exists to make sure the systems we’re building today can be trusted in production tomorrow, and to set the bar for what “production‑ready” means on this team. The engineer in this role delivers their mandate by writing production code, shaping architecture, and engineering the systems themselves — not by absorbing operational load.

You’ll be responsible for

Define and operate against SLOs. Establish meaningful SLIs and SLOs with product and engineering partners, manage error budgets, and use them as real inputs to prioritization rather than dashboards no one reads.

  • Build the observability layer. Improve metrics, logs, traces, and alerting so issues are detected early, attributed precisely, and debugged with code‑level confidence. Push instrumentation upstream into the services we own.
  • Lead incident response. Act as incident commander when needed, drive blameless post‑mortems, and turn findings into concrete engineering work that lands. Build the muscle in the team so this isn’t centralized in any one person.
  • Reduce toil through engineering. Identify repetitive operational work and eliminate it with software — automation, self‑healing behavior, better defaults, better tooling — rather than absorbing it as ongoing overhead.
  • Production Hardening. Stress‑test designs for partial failure, dependency degradation, traffic spikes, and adversarial inputs. Run capacity and performance work before incidents arise. Ensure resiliency primitives are tuned and working correctly.
  • Make change safe and fast. Improve release safety through progressive delivery, feature flags, canaries, rollbacks, and tested migrations. Help the squad ship faster and with lower blast radius.
  • Improve developer experience especially where it removes operational friction or improves change safety. Where internal tooling or platform gaps slow the team down, build or contribute the fix. Prefer leverage over heroics.
  • Partner across disciplines. Work closely with product, platform, security, compliance, and other engineering teams. Translate reliability and risk tradeoffs into language each audience can act on.
  • Raise the engineering bar. Mentor engineers, review hard designs and PRs, and shape technical standards across the squad. Lead through clarity and judgment, not authority.
We are looking for a person who has

Track record of owning services in production — not just shipping them, but being the engineer responsible for how they behave under real load and real failure.

  • Experience defining and operating against SLOs/SLIs, and using error budgets to influence engineering and product decisions.
  • Experience leading incident response and writing post‑mortems that produced durable improvements.
  • Hands‑on experience with observability tooling (metrics, structured logging, distributed tracing) and using it to diagnose nontrivial production issues.
  • Deep system design experience: distributed services, asynchronous messaging, storage tradeoffs, API design, idempotency, consistency, backpressure, and graceful degradation.
  • Significant industry experience building and operating production software systems in a high‑ownership engineering environment.
  • Comfort operating in modern cloud environments (e.g., AWS/GCP), containerized workloads, and CI/CD pipelines, and reasoning about their failure modes.
  • Demonstrated technical leadership: influencing architecture across teams, mentoring strong engineers, and making the people around you better.
  • Pragmatism. You can hold a high reliability bar while still helping a fast‑moving squad ship.
Location for this opportunity (City, Country)
  • Miami, United States
  • Opportunity of earning equity at Nu
  • Medical Insurance
  • Dental and Vision Insurance
  • Life Insurance and AD&D
  • Extended maternity and paternity leaves
  • Nucleo - Our learning platform of courses
  • NuCare - Our mental health and wellness assistance program
  • 401K
  • Saving Plans - Health Saving Account and Flexible Spending Account
  • Work‑from‑home Allowance
  • Relocation Assistance Package, if applicable.
#J-18808-Ljbffr