1

Director Site Reliability Engineering Jobs (NOW HIRING)

Director, Site Reliability Engineering

$58.25 - $77.50/hr

We're looking for a Senior Manager of Site Reliability Engineering to join our team. You'll lead a team of ~10 SREs across North America, UK, HK, and New Zealand - owning both the day-to-day ...

Site Reliability Engineer

San Francisco, CA ยท Remote

$67.25 - $89.25/hr

Director, Site Reliability Location: Remote (US) Department: Cloud Platform Engineering / SRE/Reliability Position summary The Site Reliability Engineer (SRE) owns reliability, observability, and ...

Site Reliability Engineering

Los Angeles, CA ยท On-site

$61.50 - $81.50/hr

Site Reliability Engineering (SRE) Location: Los Angeles, CA Remote position Fulltime position JD * Site Reliability Engineer * Experience in Cloud platforms (AWS, Azure, Google Cloud) and hybrid ...

... rather than direct authority. This role is accountable for building and leading the SRE team ... including hiring, performance management, coaching, and development of engineers transitioning into ...

New

Showing results 21-40

Director Site Reliability Engineering information

See salary details

$10

$63

$91

How much do director site reliability engineering jobs pay per hour?

As of Aug 8, 2026, the average hourly pay for director site reliability engineering in the United States is $63.74, according to ZipRecruiter salary data. Most workers in this role earn between $54.81 and $72.84 per hour, depending on experience, location, and employer.

How much do director site reliability engineering get paid?

Director of Site Reliability Engineering typically earns a salary ranging from $150,000 to $250,000 annually, depending on experience, company size, and location. They often oversee teams using tools like Kubernetes and Prometheus and require strong leadership and technical skills.

What is a director site reliability engineering?

A Director of Site Reliability Engineering (SRE) leads teams responsible for ensuring the availability, performance, and scalability of software systems. They define reliability best practices, drive automation, and collaborate with engineering and product teams to improve system resilience. This role requires strong leadership, technical expertise, and a focus on balancing innovation with operational stability.

What are the main challenges faced by a director site reliability engineering, and how can I prepare for them?

A Director of Site Reliability Engineering often encounters challenges such as balancing rapid feature delivery with system stability, managing complex incident responses, and fostering a culture of continuous improvement. Additionally, aligning reliability goals with business objectives and securing cross-functional buy-in can be demanding. To prepare, it is helpful to gain experience in high-scale system management, develop strong leadership and communication abilities, and cultivate a proactive approach to risk management and automation. Staying up to date with the latest SRE practices and building relationships with both engineering and business teams will also support your success in this pivotal role.

What are the key skills and qualifications needed to thrive as a director site reliability engineering?

To thrive as a Director Site Reliability Engineering, you need extensive experience in software engineering, infrastructure management, incident response, and people leadership, often supported by a degree in computer science or a related field. Familiarity with cloud platforms (such as AWS, GCP, or Azure), automation tools (Terraform, Ansible), monitoring systems (Prometheus, Datadog), and relevant certifications like CKA or AWS Solutions Architect is valued. Outstanding communication, stakeholder management, and strategic vision are key soft skills that set leaders apart in this role. These abilities ensure the reliability, scalability, and efficiency of critical systems while effectively guiding and motivating technical teams.

More about Director Site Reliability Engineering jobs
What cities are hiring for Director Site Reliability Engineering jobs? Cities with the most Director Site Reliability Engineering job openings:
What states have the most Director Site Reliability Engineering jobs? States with the most job openings for Director Site Reliability Engineering jobs include:
What job categories do people searching Director Site Reliability Engineering jobs look for? The top searched job categories for Director Site Reliability Engineering jobs are:
Infographic showing various Director Site Reliability Engineering job openings in the United States as of August 2026, with employment types broken down into 90% Full Time, and 10% Part Time. Highlights an 70% In-person, 5% Hybrid, and 25% Remote job distribution, with an average salary of $132,583 per year, or $63.7 per hour.

Director, Site Reliability Engineering

Stellar Development Foundation

San Francisco, CA โ€ข On-site

$67.25 - $89.25/hr

Full-time

Medical, Dental, Vision, Life, Retirement

Re-posted 6 days ago


Job description

Interested in working on cutting-edge blockchain technology and creating equitable access to the global financial system? Since 2014, the mission-driven team at the Stellar Development Foundation (SDF) has helped fuel the tremendous growth of the Stellar blockchain network, an open-source platform that operates at high-scale today. Developers and companies around the world build on it, and the SDF team is expanding to support the rapidly growing and changing Stellar ecosystem.
SDF is looking for a Director of Site Reliability Engineering to lead a small, high-leverage SRE team and help shape how engineering teams own, operate, and improve production services.
This is a senior engineering leadership role reporting to the CTO. You will set the vision, operating model, and culture for SRE while owning the core infrastructure services that help SDF engineering teams build, deploy, observe, and operate software with confidence.
Engineering teams at SDF own the services they build. SRE provides the frameworks, standards, shared infrastructure, tooling, observability practices, and enablement model that make strong service ownership possible across engineering.
You will be successful here if you bring strong technical judgment, pragmatic leadership, and the ability to influence through trust, clarity, and execution. SDF is a small, mission-driven foundation with a broad technical surface area, so this role requires leverage, ownership, and a bias toward solving the right problems over creating processes for its own sake.
In this role, you will:
  • Lead, coach, and develop a distributed SRE team, setting a clear vision, charter, operating model, priorities, and success measures.
  • Define and roll out a Service Ownership & Maturity Framework across engineering, with expectations that vary appropriately by service criticality.
  • Own and improve core engineering infrastructure services, including cloud foundations, Kubernetes and compute patterns, CI/CD, observability, secrets management, GitHub workflows, and infrastructure automation.
  • Help engineering teams become stronger owners and operators of their services through better standards, dashboards, runbooks, alerting, escalation paths, operational readiness, and deployment practices.
  • Make reliability, operational maturity, infrastructure health, and developer productivity more measurable through trusted metrics and practical operational intelligence.
  • Improve deployment automation, resilience, self-healing patterns, disaster recovery readiness, and service reliability based on actual impact and risk.
  • Mature incident response, escalation, postmortems, and on-call health across a geographically distributed team.
  • Build paved paths and self-service infrastructure that reduce toil, lower cognitive load, and help engineering teams move faster while strengthening ownership and reliability.
  • Partner closely with Security, Compliance, Legal, Finance, Procurement, and Corporate IT where infrastructure, access management, cloud operations, vendor review, or controls intersect with engineering.
  • Pragmatically evaluate AI-assisted and agentic workflows where they can improve infrastructure operations, service ownership, developer workflows, or toil reduction.

You have:
  • 10+ years of experience in SRE, infrastructure engineering, platform engineering, cloud infrastructure, production operations, or closely related engineering roles.
  • 5+ years of experience leading, managing, or formally developing infrastructure, SRE, platform, or reliability engineers.
  • Strong experience defining team charters, operating models, roadmaps, success measures, and engineering practices for infrastructure or reliability teams.
  • Deep technical judgment across cloud infrastructure, production operations, distributed systems, reliability tradeoffs, automation, and operational risk.
  • 3+ years of experience with modern cloud infrastructure in AWS, GCP, or similar environments.
  • 3+ years of experience with Kubernetes, container orchestration, infrastructure-as-code, declarative systems, CI/CD, and deployment safety.
  • Strong experience with observability, monitoring, alerting, logging, dashboards, SLOs/SLIs, incident response, postmortems, and on-call practices.
  • Experience helping product or application engineering teams improve service ownership, operational readiness, and production accountability.
  • A pragmatic approach to tooling: you understand when to build, buy, adapt, simplify, or retire systems based on the actual engineering problem.
  • The ability to operate effectively in a small or mid-sized engineering organization where influence comes from credibility, judgment, and outcomes rather than bureaucracy.
  • Clear executive communication skills and the ability to partner directly with a CTO and senior engineering leaders.

Bonus Points if:
  • Experience leading SRE, infrastructure, or platform work in a lean, high-agency organization.
  • Experience supporting globally distributed teams or 24/7 operational coverage.
  • Experience improving developer productivity through paved paths, self-service infrastructure, automation, and reduced toil.
  • Experience with infrastructure security fundamentals, secrets management, access controls, cloud security practices, or compliance-related infrastructure controls.
  • Experience in financial services, regulated environments, blockchain, crypto, Web3, or other high-reliability technical ecosystems.
  • Experience evaluating vendors and infrastructure platforms with skepticism, technical rigor, and cost discipline.
  • Practical experience applying AI-assisted or agentic workflows to infrastructure, reliability, operations, observability, or developer productivity.

We offer competitive pay with a base salary range for this position of $205,000 - $305,000 depending on job-related knowledge, skills, experience, and location. In addition, we offer lumen-denominated grants along with the following perks and benefits:
USA Benefits/Perks:
  • Competitive health, dental & vision coverage with most plans covered at 100% for the employee + any dependents
  • Flexible time off + 15 company holidays including a company-wide holiday break
  • Generous paid parental leave for all parents, plus paid pregnancy disability leave for birthing parents
  • Gym reimbursement ($80 per month)
  • Life & ADD (up to $50K)
  • Short & Long term disability
  • 401K with 4% match
  • Health & Dependent Care FSA Accounts
  • Commuter benefits with $250/month employer contribution
  • Health Savings Account (HSA) with monthly employer contribution
  • Family building benefits through Kindbody
  • Wellbeing benefits (One Medical, Rightway, Headspace)
  • L&D budget of $1,500/year
  • Daily lunch and snacks in office
  • Company retreats

#LI-Hybrid
About Stellar
Stellar is more than a blockchain. Powered by a decentralized, fast, scalable, and uniquely sustainable network made for financial products and services and a thriving and passionate ecosystem that includes a non-profit organization driven by a mission, Stellar is paving the path to unlock the world's economic potential through blockchain technology. Built with speed and low costs in mind, the Stellar network provides builders and financial institutions worldwide a platform to issue assets, and to send and convert currencies in real time creating real world utility. Founded in 2014, the Stellar Development Foundation (SDF) supports the continued development and growth of the Stellar network and also serves the ecosystem of NGOs, corporations, universities, small businesses, governments, and solo entrepreneurs building on the Stellar network through tooling, funding and strategic collaborations. Together, Stellar is where blockchain meets the real world.
About the Stellar Development Foundation
The Stellar Development Foundation (SDF) is a non-profit organization focused on working with and supporting change-makers to create equitable access to the global financial system through blockchain technology. SDF provides grants, investments, funding, and other awards to builders and organizations. SDF also develops resources and tooling on the Stellar network to help unlock real world utility. As a nonprofit foundation, SDF puts the health of the Stellar network and the Stellar ecosystem and its mission above all else.
We look forward to hearing from you!
Privacy Policy
By submitting your application, you are agreeing to our use and processing of your data in accordance with our Privacy Policy.
SDF is committed to diversity in its workforce and is proud to be an equal opportunity employer. SDF does not make hiring or employment decisions on the basis of race, color, religion, creed, gender, national origin, age, disability, veteran status, marital status, pregnancy, sex, gender expression or identity, sexual orientation, citizenship, or any other basis protected by applicable local, state or federal law.