1

Director Site Reliability Engineering Jobs in Texas

SRE Lead

Plano, TX ยท On-site

$54.50 - $72.50/hr

Apply Site Reliability Engineering (SRE) principles to enhance system reliability and performance ... com Direct: 4705239688 Led by 25+ Years of Industry Experience E-Verify ยฎ is a registered ...

Site Reliability Engineer III

Plano, TX ยท On-site

$53.25 - $70.75/hr

Formal training or certification on site reliability engineering concepts and 3+ years applied experience * Exposure to or hands-on experience in supporting SRE practices for Data management ...

Site Reliability Engineer III

Plano, TX ยท On-site

$53.25 - $70.75/hr

Formal training or certification on site reliability engineering concepts and 3+ years applied experience * Exposure to or hands-on experience in supporting SRE practices for Data management ...

Site Reliability Engineer III

Plano, TX ยท On-site

$53.25 - $70.75/hr

Formal training or certification on site reliability engineering concepts and 3+ years applied experience * Exposure to or hands-on experience in supporting SRE practices for Data management ...

Site Reliability Engineer III

Plano, TX ยท On-site

$53.25 - $70.75/hr

Formal training or certification on site reliability engineering concepts and 3+ years applied experience * Exposure to or hands-on experience in supporting SRE practices for Data management ...

Site Reliability Engineer III

Plano, TX ยท On-site

$53.25 - $70.75/hr

Formal training or certification on site reliability engineering concepts and 3+ years applied experience * Exposure to or hands-on experience in supporting SRE practices for Data management ...

Showing results 41-60

Director Site Reliability Engineering information

See Texas salary details

$10

$59

$85

How much do director site reliability engineering jobs pay per hour?

As of Sep 6, 2026, the average hourly pay for director site reliability engineering in Texas is $59.39, according to ZipRecruiter salary data. Most workers in this role earn between $51.06 and $67.84 per hour, depending on experience, location, and employer.

What is a director site reliability engineering?

A Director of Site Reliability Engineering (SRE) leads teams responsible for ensuring the availability, performance, and scalability of software systems. They define reliability best practices, drive automation, and collaborate with engineering and product teams to improve system resilience. This role requires strong leadership, technical expertise, and a focus on balancing innovation with operational stability.

What are the key skills and qualifications needed to thrive as a director site reliability engineering?

To thrive as a Director Site Reliability Engineering, you need extensive experience in software engineering, infrastructure management, incident response, and people leadership, often supported by a degree in computer science or a related field. Familiarity with cloud platforms (such as AWS, GCP, or Azure), automation tools (Terraform, Ansible), monitoring systems (Prometheus, Datadog), and relevant certifications like CKA or AWS Solutions Architect is valued. Outstanding communication, stakeholder management, and strategic vision are key soft skills that set leaders apart in this role. These abilities ensure the reliability, scalability, and efficiency of critical systems while effectively guiding and motivating technical teams.

What are the main challenges faced by a director site reliability engineering, and how can I prepare for them?

A Director of Site Reliability Engineering often encounters challenges such as balancing rapid feature delivery with system stability, managing complex incident responses, and fostering a culture of continuous improvement. Additionally, aligning reliability goals with business objectives and securing cross-functional buy-in can be demanding. To prepare, it is helpful to gain experience in high-scale system management, develop strong leadership and communication abilities, and cultivate a proactive approach to risk management and automation. Staying up to date with the latest SRE practices and building relationships with both engineering and business teams will also support your success in this pivotal role.

How much do Director Site Reliability Engineers get paid?

Director Site Reliability Engineers typically earn between $130,000 and $200,000 annually, depending on experience, location, and company size. They often oversee large teams, require advanced skills in cloud platforms and automation tools, and may receive bonuses or stock options as part of compensation.

What are popular job titles related to Director Site Reliability Engineering jobs in Texas?

For Director Site Reliability Engineering jobs in Texas, the most frequently searched job titles are:

What cities in Texas are hiring for Director Site Reliability Engineering jobs?

Cities in Texas with the most Director Site Reliability Engineering job openings:

Infographic showing various Director Site Reliability Engineering job openings in Texas as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 13% Part Time, 1% Temporary, 2% Contract, and 1% Nights. Highlights an 93% Physical, 2% Hybrid, and 5% Remote job distribution, with an average salary of $123,522 per year, or $59.4 per hour.

Director - Application Site-Reliability Engineering

Caris MPI, Inc.

Irving, TX โ€ข On-site

$160 - $210/hr

Other

Posted 8 days ago


Job description

At Caris, we understand that cancer is an ugly wordโ€”a word no one wants to hear, but one that connects us all. Thatโ€™s why weโ€™re not just transforming cancer careโ€”weโ€™re changing lives.We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: โ€œWhat would I do if this patient were my mom?โ€ That question drives everything we do.But our mission doesnโ€™t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcareโ€”driven by innovation, compassion, and purpose.Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.Position SummaryCaris Life Sciences is one of the largest precision-oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double-digit rates. Behind every case is a matched molecular, imaging, and clinical-outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it reliable is this role's charter.Reporting to the Corporate Vice President for Clinical Software Products, the Director owns application reliability and production operations for the clinical software portfolio: production support, incident response, on-call operations, and SLO management for applications under SOX financial controls and FDA regulatory requirements. The Director hires and develops the team, sets standards and selects tooling, establishes production-access governance and segregation-of-duties controls with the information-security, quality, and infrastructure organizations, and participates directly in incident response. The role carries wide latitude to shape how reliability engineering is done here.Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning incident diagnostics, runbook authoring, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance.The infrastructure organization owns the platform and observability runtime; this role owns application-layer reliability on top of it. The operating model is automation-first: recurring manual work is engineered away rather than staffed, and operational and compliance evidence is produced by pipelines rather than assembled by hand.Location: Irving, Texas (Dallasโ€“Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.Job ResponsibilitiesOwn the production-support model for the clinical application portfolio: the on-call rotation, escalation procedures, and incident-response playbooks.Establish and audit production-access and segregation-of-duties controls with engineering, information-security, quality, and infrastructure partners; keep the evidence audit-ready for SOX ITGCs and applicable FDA requirements, including access grants, role changes, and privileged-action logs.Define clinical SLOs, error budgets, and availability targets with product and engineering leadership; track attainment and operational-health metrics such as mean time to detect, mean time to recover, and on-call burden; intervene while an error budget is burning, not after it is spent.Lead incident response for high-severity production events as incident commander or senior technical responder; coordinate cross-functional teams, run post-incident reviews, and drive systemic remediation to closure.Develop the runbook library and approved operational automation; ensure engineers can execute standard interventions safely within documented procedures.Drive the automation-first operating model: convert recurring manual interventions into reviewed automation, measure and reduce toil, and favor self-service tooling over ticket-driven request work.Coordinate production deployments with engineering teams, verify deployment health, and own rollback decisions.Advance automated generation of change and deployment evidence in CI/CD pipelines: deployment records, approval trails, and change documentation as an audit-ready by-product of release.Hire and develop the App-SRE team; set performance expectations, on-call responsibilities, and career growth frameworks.Own application-layer observability alongside the infrastructure and observability platform teams; keep production dashboards, alerting thresholds, and SLO monitors accurate and actionable.Represent App-SRE in engineering leadership forums, operational reviews, and compliance audits; translate operational health and risk into clear executive communication.Run the function AI-first: make AI-assisted practice the team's daily norm, from runbook automation to incident analysis and operational tooling.Required QualificationsBachelor's degree in Computer Science, Software Engineering, Information Systems, or a closely related technical field, or equivalent practical experience.10+ years of professional experience in SRE, DevOps, platform engineering, or production operations.4+ years of direct experience in a people-management or team-lead role within an SRE or production-operations function.Hands-on experience leading incident response for Tier 1 or business-critical production systems, including serving as incident commander or senior technical responder.Experience defining and implementing SLOs, error budgets, and associated alerting and on-call workflows in a production environment.Experience building or significantly maturing a production-support, on-call, or SRE function, including runbook development and on-call-rotation design.Experience operating production systems under a formal regulatory or financial-controls framework, such as CAP/CLIA, FDA regulations, SOX ITGCs, HIPAA, or equivalent, including producing documentation that holds up in audit.A record of applying AI-assisted practice to operations or engineering work, personally or through a team.Preferred QualificationsDomain experience in clinical diagnostics, laboratory information systems, molecular pathology, or digital health software.Direct experience supporting SOX ITGC audit cycles or CAP/CLIA laboratory inspections, including evidence gathering for access-control, change-management, and monitoring controls.Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.Experience with deployment pipelines, release-management workflows, and rollback procedures in a continuous-delivery environment.Ability to operate as a player-coach, contributing directly to technical work while building and leading a team.Track record of reducing operational toil through automation programs in an SRE or production-operations organization.Experience presenting operational strategy and risk posture to senior or executive audiences.Physical DemandsAbility to sit, stand, and work at a computer for extended periods.TrainingAll job-specific, safety, and compliance training is assigned based on the job functions associated with this employee.OtherThis role serves in the senior escalation tier of the production on-call rotation, with after-hours response to high-severity incidents as incident commander or senior escalation point. Periodic travel may be required to support business needs, team on-sites, and leadership reviews.Conditions of Employment: Individual must successfully complete pre-employment process, which includes criminal background check, drug screening, credit check ( applicable for certain positions) and reference verification.This job description reflects managementโ€™s assignment of essential functions. Nothing in this job description restricts managementโ€™s right to assign or reassign duties and responsibilities to this job at any time.Caris Life Sciences is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, gender identity, sexual orientation, age, status as a protected veteran, among other things, or status as a qualified individual with disability. #J-18808-Ljbffr