1

Reliability Engineer Jobs in Austin, TX (NOW HIRING)

Site Reliability Engineer (SRE)

Austin, TX ยท On-site

$56.50 - $75/hr

Site Reliability Engineer (SRE) Location: Austin, TX Job Type: Full Time Technical Skills: * 6+ years of professional engineering experience developing, managing, or supporting distributed systems ...

Site Reliability Engineer (SRE)

Austin, TX ยท On-site

$56.50 - $75/hr

Site Reliability Engineer (SRE) Location: Austin, TX Job Type: Full Time Job Summary - Seasoned Site Reliability Engineer (SRE) with 7+ years of experience in supporting complex, large-scale ...

SRE Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

InterSources Inc is currently seeking a highly skilled SRE hands-on Lead Engineer with solid experience to help lead transformational initiatives within IT operations. In this role, you will design ...

Title: Site Reliability Engineer SRE - ML platform Location: Austin, TX OR Sunnyvale, CA Type: FTE Salary/Rate : $140K Title: Site Reliability Engineer SRE - ML platform Responsibilities

Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

Site Reliability Engineer SRE - ML platform Location: Austin, TX OR Sunnyvale, CA Title: Site Reliability Engineer SRE - ML platform Responsibilities - * Continuous Deployment using GitHub Actions ...

Principal Reliability Engineer

Austin, TX

$101K - $127K/yr

We are looking for a Principal Reliability Engineer to join our team located in Tucson, AZ. What You Will Do: * Develop and implement reliability strategy across multiple product phases including ...

SRE/DevOps Engineer - GenAI

Austin, TX ยท On-site

$56.50 - $75/hr

Core SRE / DevOps Skills * Strong experience in DevOps / SRE roles supporting production systems * Hands-on expertise with AWS, including EKS * Solid experience with Docker, Kubernetes, and Helm

Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

Future Secure AI is building innovative solutions at the forefront of AI technology, seeking a Site Reliability Engineer to design, build, and operate the platforms that power AI Co-Workers. The role ...

Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

They are seeking a Site Reliability Engineer to design, build, and operate platforms that support AI Co-Workers, ensuring reliability and efficiency throughout the software lifecycle.

Sr. SRE - Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

As a Senior Site Reliability Engineer, you will serve as a technical leader responsible for advancing platform reliability, resiliency, and operational excellence across complex distributed systems.

Reliability Engineer

Austin, TX

$101K - $127K/yr

Duties: Perform semiconductor engineering design, development, and testing for firm products and ... Apply quality and reliability procedures and requirements, as well as to other core team processes ...

Reliability Engineer

Austin, TX ยท On-site

$105 - $125/hr

Duties Perform semiconductor engineering design, development, and testing for firm products and ... Apply quality and reliability procedures and requirements, as well as to other core team processes ...

Quality & Reliability Engineer

Austin, TX ยท On-site

$101K - $127K/yr

We're hiring a Quality & Reliability Engineer to own both the quality of our supply chain and manufacturing operations and the long-term dependability of the P3 system in the field. You'll qualify ...

Quality & Reliability Engineer

Austin, TX ยท On-site

$101K - $127K/yr

We're hiring a Quality & Reliability Engineer to own both the quality of our supply chain and manufacturing operations and the long-term dependability of the P3 system in the field. You'll qualify ...

Quality & Reliability Engineer

Austin, TX

$101K - $127K/yr

We're hiring a Quality & Reliability Engineer to own both the quality of our supply chain and manufacturing operations and the long-term dependability of the P3 system in the field. You'll qualify ...

Reliability Engineer

Austin, TX ยท On-site

$101K - $127K/yr

Duties: Perform semiconductor engineering design, development, and testing for firm products and ... Apply quality and reliability procedures and requirements, as well as to other core team processes ...

next page

Showing results 1-20

Reliability Engineer information

See Austin, TX salary details

$60.4K

$116.9K

$139.7K

How much do reliability engineer jobs pay per year?

As of Aug 24, 2026, the average yearly pay for reliability engineer in Austin, TX is $116,908.00, according to ZipRecruiter salary data. Most workers in this role earn between $101,600.00 and $127,800.00 per year, depending on experience, location, and employer.

What is a reliability engineer?

Reliability Engineers are professionals responsible for ensuring that systems, equipment, or processes function consistently and efficiently over time. They analyze data, identify potential points of failure, and develop maintenance strategies to improve system reliability and minimize downtime. Their work spans various industries, including manufacturing, energy, and technology, and often involves collaborating with design, operations, and maintenance teams. By implementing reliability-centered maintenance and predictive analysis, they help organizations save costs and increase safety.

What does a reliability engineer do?

As a reliability engineer, your duties are to test and evaluate the manufacturing of products and components and ensure that the procedures are efficient and do not lead to abnormally high maintenance or operational costs. Your other responsibilities are to find solutions to product reliability risks. You may manage risk in a supply chain, develop loss prevention strategies, and track the entire lifecycle of product development, from building prototypes to moving a product into full-scale production. You analyze information from department heads and recommend strategies to reduce risk and ensure that the product works reliably.

What are the key skills and qualifications needed to thrive as a reliability engineer, and why are they important?

To thrive as a Reliability Engineer, you need a solid background in engineering principles, failure analysis, and reliability modeling, typically with a degree in engineering or a related field. Familiarity with tools such as FMEA, Root Cause Analysis (RCA), reliability-centered maintenance (RCM) software, and certifications like Certified Reliability Engineer (CRE) are highly valued. Strong problem-solving abilities, attention to detail, and effective communication are crucial soft skills in this role. These skills ensure systems are dependable, downtime is minimized, and organizational performance and safety are optimized.

What are some typical challenges reliability engineers face when implementing preventive maintenance strategies?

Reliability Engineers often encounter challenges such as balancing preventive maintenance schedules with production demands, ensuring buy-in from operations teams, and accurately predicting equipment failures. They must analyze large sets of historical data to identify trends and root causes, which can be complex in facilities with diverse machinery. Collaboration with maintenance, operations, and engineering teams is essential to develop effective strategies that minimize downtime while optimizing resources.

What is the difference between Reliability Engineer vs Maintenance Engineer?

AspectReliability EngineerMaintenance Engineer
CredentialsTypically requires engineering degree, certifications in reliability or asset managementOften requires engineering or technical diploma, certifications in maintenance or equipment repair
Work EnvironmentFocuses on analysis, design, and improvement of systems for reliabilityHands-on maintenance, repair, and troubleshooting of equipment
Industry UsageCommon in manufacturing, energy, aerospace, and industrial sectorsPrevalent in manufacturing, facilities, and industrial plants

Reliability Engineers focus on designing and improving systems to prevent failures, using data analysis and modeling. Maintenance Engineers perform hands-on repairs and upkeep of equipment to ensure operational continuity. While both roles aim to optimize equipment performance, Reliability Engineers work proactively on system reliability, whereas Maintenance Engineers handle reactive and scheduled maintenance tasks.

Are reliability engineers in demand?

Reliability engineers are in high demand across industries such as manufacturing, energy, and aerospace due to their role in improving system performance and reducing downtime. Employers seek professionals with skills in data analysis, failure modes, and maintenance strategies, often requiring certifications like Certified Reliability Engineer (CRE). The job outlook is positive, with steady growth expected as companies prioritize operational efficiency and risk management.

How much do reliability engineers get paid?

Reliability engineers typically earn a median annual salary ranging from $70,000 to $110,000, depending on experience, location, and industry. Senior or specialized reliability engineers with certifications and advanced skills can earn higher salaries, often exceeding $120,000 annually.

What are the most commonly searched types of Reliability Engineer jobs in Austin, TX?

The most popular types of Reliability Engineer jobs in Austin, TX are:

What are popular job titles related to Reliability Engineer jobs in Austin, TX?

For Reliability Engineer jobs in Austin, TX, the most frequently searched job titles are:

What job categories do people searching Reliability Engineer jobs in Austin, TX look for?

The top searched job categories for Reliability Engineer jobs in Austin, TX are:

What cities near Austin, TX are hiring for Reliability Engineer jobs?

Cities near Austin, TX with the most Reliability Engineer job openings:

Infographic showing various Reliability Engineer job openings in Austin, TX as of August 2026, with employment types broken down into 91% Full Time, and 9% Contract. Highlights an 91% In-person, and 9% Remote job distribution, with an average salary of $116,908 per year, or $56.2 per hour.

Senior Systems Reliability Engineer

Electric Reliability Council of Texas

Taylor, TX โ€ข On-site

$120 - $170/hr

Other

Posted 6 days ago


Job description

## Senior Systems Reliability EngineerApplylocations: Taylor, TXtime type: Full timeposted on: Posted Todayjob requisition id: R2378At ERCOT, our diverse and dynamic work environment provides a platform on which employees can work together to build the future of the Texas power grid and wholesale market utilizing the latest technologies and resources. We encourage you to join our talented, dedicated workforce to develop world-class solutions for today and tomorrowโ€™s energy challenges while learning new skills and growing your career.ERCOT is committed to fostering inclusion at all levels of our company. It is the cornerstone of our corporate values of accountability, leadership, innovation, trust, and expertise. We know that individuals with a wide variety of talents, ideas, and experiences propel the innovation that drives our success. An inclusive and diverse workforce strengthens us and allows for a collaborative environment to solve the challenges that face our industry today and in the future.**JOB SUMMARY**The Senior Systems Reliability Engineer applies software engineering discipline to reliability problems โ€” designing, building, and operating the systems that make production software measurable, scalable, and self-healing. This role treats operational challenges as engineering problems: when a process is manual, it gets automated; when a failure mode is unknown, it gets instrumented; when a system degrades, the degradation is understood before it recurs.At this level, the specialist owns SLO and error budget frameworks for assigned systems, architects the observability stack that the team relies on, leads engineering-driven incident response, and holds NERC/CIP compliance responsibility for assigned systems. This role partners directly with Software Engineers as a technical peer โ€” participating in design reviews, influencing architecture decisions for reliability, and building the production readiness standards that govern how software ships. Advancement to Lead is based on demonstrated ability to define reliability engineering standards at the platform level, influencing practice across multiple teams and portfolios.**JOB DUTIES*** Performs complex reliability engineering work autonomously; recognized subject matter expert within the team and adjacent teams.* Designs and builds production software systems, reliability tooling, and automation frameworks; treats operational problems as engineering problems to be solved through code.* Owns SLO governance, error budget management, and observability architecture for assigned systems; leads engineering-driven incident response including failover scenarios.* Holds NERC/CIP compliance responsibility for assigned systems; formally mentors less experienced specialists; may coordinate team delivery and on-call activities.**ADDITIONAL JOB DUTIES****Core Expectations**The following expectations apply at all Systems Reliability Specialist levels. Scope and independence expand with each level.* Engineer reliability solutions: when a process is manual and repeatable, automate it; when a failure mode is opaque, instrument it; when a system is fragile, redesign the failure boundary.* Define and own SLIs and SLOs for assigned systems; treat error budgets as a shared engineering contract with development teams, not an operations metric.* Respond to production incidents as an engineer: form a hypothesis, isolate the failure, resolve it, and close the loop with a post-mortem that addresses root cause.* Instrument systems so that on-call responders have sufficient telemetry to diagnose and act without tribal knowledge.* Participate in 24/7 on-call rotation; treat every alert as signal โ€” either actionable or worth eliminating.* Write production-quality code: reliability tooling, automation frameworks, and operational software are held to the same engineering standards as application code.* Partner with development teams as a peer in design reviews; reliability is designed in, not bolted on after deployment.**Reliability Engineering**Senior specialists design and build the engineering systems that make production software reliable. This is software engineering applied to operational problems โ€” the output is code, frameworks, and automated systems, not tickets and runbooks alone.* Design, build, and maintain reliability tooling: automated remediation systems, self-healing infrastructure components, and operational software that reduces human intervention in production.* Own SLO and error budget definitions for assigned systems; review error budget consumption with development teams and drive engineering decisions based on budget status.* Architect and implement chaos engineering programs: define failure injection scenarios, automate resilience tests, and validate recovery behavior against defined SLOs.* Build and maintain CI/CD reliability gates: automated canary analysis, progressive delivery validation, and rollback triggers based on SLI thresholds.* Design capacity planning models for assigned systems; build tooling to project resource needs and surface capacity risks before they affect availability.* Contribute to production readiness reviews: define and enforce the engineering criteria that a system must meet before it ships to production.* Reduce operational toil through engineering: measure toil, track reduction targets, and build the automation that eliminates it.**Observability & Instrumentation**Observability is an engineering discipline. Senior specialists design and build the telemetry systems that make production behavior understandable โ€” not just monitored.* Architect MLTP (Metrics, Logs, Traces, Profiling) observability solutions using the Grafana LGTM stack (Loki, Grafana, Tempo, Mimir), Dynatrace APM, Splunk, and Datadog.* Define and enforce instrumentation standards: structured logging schemas, metric naming conventions, trace context propagation, and continuous profiling configuration for assigned systems.* Build distributed tracing coverage across service boundaries; identify and close observability gaps that produce blind spots during incidents.* Design SLI instrumentation: translate user-facing reliability requirements into specific, measurable signals that accurately represent system health from the user's perspective.* Build and maintain alerting frameworks: alerts must be actionable, calibrated to SLO burn rate, and free of noise; own alert quality as an engineering output.* Correlate application performance data โ€” JVM heap behavior, GC pressure, thread contention โ€” with infrastructure events to enable root cause analysis across layers.**Incident Response & Problem Management**Incident response at this level is an engineering activity. Senior specialists lead the technical response to high-severity events, own the post-mortem process, and drive the engineering work that prevents recurrence.* Lead high-severity incident response for assigned systems, including dual-datacenter failover execution; own the technical resolution from detection through remediation.* Apply structured root cause analysis: distinguish symptoms from causes, identify contributing factors across system layers, and drive remediation that addresses root cause rather than surface behavior.* Author post-mortems that produce actionable engineering work items โ€” not process improvements alone; track remediation to completion and validate effectiveness.* Diagnose complex cross-layer failures: Java/JVM application failures, distributed system race conditions, database connection pool exhaustion, messaging system backpressure, and cross-datacenter synchronization issues.* Build and maintain incident response runbooks as engineering artifacts: automated where feasible, version-controlled, and validated during chaos engineering exercises.* Participate in blameless post-mortem facilitation; model the engineering culture that treats incidents as system failures, not human failures.**Java Application Reliability**The primary application platform is Java/Spring Boot. Senior specialists are expected to operate at the intersection of application engineering and reliability โ€” understanding the runtime deeply enough to diagnose, tune, and improve production behavior.* Diagnose and resolve Java application performance problems in production: heap memory pressure, garbage collection tuning, thread pool exhaustion, connection leak detection, and class loading anomalies.* Perform JVM performance analysis using heap dumps, thread dumps, and continuous profiling; translate findings into engineering recommendations for development teams.* Instrument Spring Boot applications with production-grade observability: Micrometer metrics, structured logging with correlation IDs, and distributed trace integration.* Diagnose failures across the Java application stack: Spring Boot service behavior, PostgreSQL and Oracle query performance, Kafka and ActiveMQ messaging reliability, and REST/SOAP API integration failures.* Contribute to Java application design reviews with a reliability lens: identify failure modes, single points of failure, and observability gaps before code ships to production.**Platform & Infrastructure Engineering**Senior specialists build and maintain the platform engineering components that reliability depends on โ€” container orchestration, infrastructure automation, deployment tooling, and environment governance.* Design and operate Kubernetes and OpenShift workloads for reliability: resource quotas, pod disruption budgets, horizontal pod autoscaling, and liveness and readiness probe engineering.* Build infrastructure-as-code for reliability infrastructure: Terraform modules, Ansible/AAP playbooks, and Azure Resource Manager templates that are tested, version-controlled, and peer-reviewed.* Own dual-datacenter reliability architecture for assigned systems: synchronization validation, automated failover triggering, traffic management, and recovery time objective verification.* Design and automate environment promotion pipelines: ensure that configuration, secrets, and infrastructure state are consistent and validated across development, test, staging, and production.* Build and maintain automated patch compliance workflows; integrate CVE remediation into CI/CD pipelines rather than treating it as a manual operational process.**NERC/CIP Compliance**NERC/CIP compliance for assigned systems is an engineering responsibility at this level โ€” not a documentation exercise. Senior specialists implement controls through code and automation wherever possible.* Own NERC/CIP compliance for assigned systems: interpret applicable reliability standards, implement required controls, maintain evidence documentation, and prepare for regulatory audit.* Engineer compliance controls into the platform where possible: automated hardening scripts, configuration drift detection, access control validation, and audit log integrity verification.* Maintain currency on applicable NERC/CIP standards and ERCOT-specific regulatory requirements; escalate emerging compliance risks to the Lead or Manager.* Participate in regulatory audit preparation: produce control evidence, respond to auditor inquiries, and coordinate with compliance stakeholders on findings remediation.**Technical Leadership & Mentoring*** Hold formal mentoring responsibility for Systems Reliability Specialist I and II team members: structured coaching on SRE practices, code review for reliability tooling, and career development conversations.* Serve as the recognized technical authority on reliability engineering and Java application operations for the team; adjacent teams and development engineers seek out this specialist for guidance.* Lead design reviews for systems within the team's scope; identify reliability risks and observability gaps before systems reach production.* Set engineering standards for the team: post-mortem quality, observability instrumentation, chaos engineering practices, and on-call readiness.* Contribute to the broader engineering organization: internal technical talks, SRE practice documentation, and shared tooling that other teams can adopt.**EXPERIENCE*** Minimum 5 years of progressive experience in systems reliability, software with an SRE focus, or a closely related discipline.* Demonstrated experience building and operating reliability engineering systems in production: SLO frameworks, observability platforms, chaos engineering programs, and automated remediation tooling.* Strong software engineering fundamentals: proficiency in Python and Java with experience writing production-quality reliability tooling and automation.* Deep Java/Spring Boot and JVM performance expertise: heap analysis, GC tuning, thread profiling, and application instrumentation.* Expert knowledge of MLTP observability tooling: Grafana LGTM stack (Loki, Grafana, Tempo, Mimir), Dynatrace, and Splunk.* Experience with Kubernetes and OpenShift: workload design, autoscaling, pod reliability, and container networking troubleshooting.* Experience with infrastructure-as-code: Terraform, Ansible/AAP, or equivalent.* Experience leading high-severity incident response and driving blameless post-mortem programs.* Experience with dual-datacenter or hybrid cloud reliability architecture preferred.* NERC/CIP compliance experience: control implementation, audit preparation, and regulatory engagement preferred.**EDUCATION*** Bachelor's Degree: Computer Science, Software Engineering, MIS, or related field (Required)* Master's Degree: Computer Science, Software Engineering, or related field (Preferred)* A combination of education and experience that provides equivalent knowledge to a major in such fields is required.**CERTIFICATION*** Azure โ€” Preferred* Certified Kubernetes Administrator (CKA) โ€” Preferred* ITIL Foundation or Managing Professional โ€” Preferred**WORK LOCATION: Taylor, TX Hybrid 2 days per week.**The foregoing description reflects the minimum qualifications and the essential functions of the position that must be performed proficiently with or without reasonable accommodation for individuals with disabilities. It is not an exhaustive list of the duties expected to be performed, and management may, at its discretion, revise or require that other or different tasks be performed as assigned. This job description is not intended to create a contract of employment with ERCOT. Both ERCOT and the employee may exercise their employment-at-will rights at any time. #LI-DN #J-18808-Ljbffr