1

Software Reliability Engineer Jobs (NOW HIRING)

Reliability Engineer

New York, NY · On-site

$165K - $250K/yr

Reliability Engineer Location NY New York United States Business Investment Management Function ... Software development of systems, services, tools and libraries. * Improve all aspects of software ...

As a reliability expert, you will be at the forefront of maintaining and enhancing the stability ... You will work closely with cross-functional teams, including software engineers, product managers ...

Job Overview This role leads the design, development, testing, implementation, and operation of secure, scalable, resilient, and highly available software platforms using Site Reliability Engineering ...

Senior Software Engineer, Site Reliability

OR · On-site +1

$57 - $75.75/hr

As a Senior Software Engineer focused on Site Reliability Tooling , your work will directly impact the success of the SRE team and all of Upstart. Your expertise will inform the team's direction, and ...

Showing results 21-40

Software Reliability Engineer information

See salary details

$39

$67

$88

How much do software reliability engineer jobs pay per hour?

As of Aug 13, 2026, the average hourly pay for software reliability engineer in the United States is $67.07, according to ZipRecruiter salary data. Most workers in this role earn between $59.13 and $74.52 per hour, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as a software reliability engineer, and why are they important?

To thrive as a Software Reliability Engineer, you need a strong background in software development, system architecture, and incident response, often supported by a degree in computer science or related field. Familiarity with monitoring tools (like Prometheus), cloud platforms (AWS, GCP), automation frameworks, and certifications such as AWS Certified DevOps Engineer are highly valuable. Excellent problem-solving, collaboration, and communication skills help you coordinate effectively during high-pressure situations and with cross-functional teams. These abilities are crucial for maintaining system uptime, quickly resolving outages, and ensuring the overall reliability of critical software services.

What is a software reliability engineer?

Software Reliability Engineers (SREs) are IT professionals who focus on ensuring that software systems are reliable, scalable, and maintain high availability. They work at the intersection of software development and IT operations, often automating processes, monitoring system performance, and responding to incidents. SREs use engineering principles to solve operational problems, aiming to reduce downtime and improve user experience. Their responsibilities can include building tools, managing infrastructure, and collaborating with development teams to implement best practices for reliability.

How does a software reliability engineer typically interact with development and operations teams to improve system stability?

Software Reliability Engineers (SREs) work closely with both development and operations teams to ensure that systems are reliable, scalable, and maintainable. They often participate in design reviews, provide input on architectural decisions, and help define service-level objectives. SREs also collaborate with developers to automate deployment processes and create monitoring solutions, and they partner with operations staff to manage incident response and root cause analysis. This collaborative environment enables them to proactively identify potential issues and drive cross-functional improvements.

What is the difference between Software Reliability Engineer vs Software Test Engineer?

AspectSoftware Reliability EngineerSoftware Test Engineer
Primary FocusEnsuring software reliability, stability, and performance over timeDesigning and executing tests to identify bugs and verify functionality
Skills & CertificationsKnowledge of reliability engineering, scripting, monitoring toolsTesting methodologies, automation tools, scripting
Work EnvironmentCollaborates with development and operations teams, often in DevOpsWorks primarily in QA/testing teams, often in dedicated testing phases
Industry UsageCommon in software companies focusing on product stabilityWidely used in software development and QA departments

The main difference is that Software Reliability Engineers focus on maintaining long-term software stability and performance, while Software Test Engineers concentrate on identifying bugs through testing. Both roles require technical skills and often collaborate, but their core objectives differ: reliability versus defect detection.

More about Software Reliability Engineer jobs
What cities are hiring for Software Reliability Engineer jobs? Cities with the most Software Reliability Engineer job openings:
Who are the top companies hiring for Software Reliability Engineer jobs? The top employers for Software Reliability Engineer jobs are:
Infographic showing various Software Reliability Engineer job openings in the United States as of August 2026, with employment types broken down into 25% Internship, and 75% Full Time. Highlights an 100% In-person job distribution, with an average salary of $139,500 per year, or $67.1 per hour.

Software Engineer, Reliability (Avionics / Compute Systems)

Cowboy Space Corporation

Seattle, WA • On-site

Full-time

Posted 22 days ago


Job description

Job Summary:
Cowboy Space Corp. is building the infrastructure to power and connect the orbital economy. As a Software Engineer, Reliability, you will ensure the software powering avionics and high-performance compute systems operates flawlessly in orbit, taking ownership of reliability from concept through flight and building automated validation frameworks.
Responsibilities:
• Own software reliability for avionics and compute systems across the full lifecycle—from concept, design, radiation hardening, and integration to launch and on-orbit operations
• Design, develop, and maintain automated test and validation frameworks for embedded and high-performance compute systems
• Develop software to validate performance, stability, and fault tolerance of CPUs, GPUs, and custom accelerators under mission-representative conditions
• Execute system-level testing including hardware-in-the-loop (HIL), environmental, and stress testing to identify failure modes
• Characterize system performance (latency, throughput, resource utilization) and drive optimization across hardware/software boundaries
• Investigate and resolve anomalies through deep root cause analysis, using data-driven approaches to improve system reliability
• Define reliability metrics, test plans, and acceptance criteria for flight-critical systems
• Work closely with electrical, avionics, and software teams to ensure robust integration of compute platforms into spacecraft systems
• Build internal tools to support rapid debugging, telemetry analysis, and system observability
• Support vehicle and payload bring-up, integration, and pre-flight validation efforts
• Support and participate in radiation testing campaigns of avionics and compute parts
Qualifications:
Required:
• Bachelor’s degree in Computer Science, Electrical Engineering, or a related field
• 4+ years of experience in software engineering, reliability engineering, or systems validation
• Strong programming skills in Python, C++, or similar languages
• Experience working with distributed systems, embedded systems, or high-performance compute environments
• Familiarity with debugging complex system-level issues across hardware and software
• Experience building automated test frameworks or validation infrastructure
Preferred:
• Experience with avionics systems, spacecraft hardware, or embedded compute platforms
• Familiarity with GPU/accelerator-based systems and performance validation (e.g., CUDA, parallel computing)
• Experience with hardware-in-the-loop (HIL) testing and system integration
• Strong background in performance analysis, profiling, and optimization
• Experience with observability tools, logging, and telemetry systems
• Understanding of fault tolerance, redundancy, and high-reliability system design
• Experience operating in fast-paced, high-ownership environments (e.g., aerospace, defense, advanced hardware startups)
Company:
Cowboy Space Corporation is building a power grid in outer space for artificial intelligence. Founded in 2024, the company is headquartered in San Carlos, USA, with a team of 51-200 employees. The company is currently Growth Stage.