1

Site Reliability Engineer Manager Jobs in Maryland

Site Reliability Engineer

Frederick, MD · Hybrid

$56.75 - $75.25/hr

With experts in biomedical science, software engineering, and program management, we focus on ... Transportation Reimbursement Account (TRN) The Site Reliability Engineer role centers on ...

Site Reliability Engineer

Frederick, MD · Hybrid

$56.75 - $75.25/hr

With experts in biomedical science, software engineering, and program management, we focus on ... Transportation Reimbursement Account (TRN) The Site Reliability Engineer role centers on ...

Site Reliability Engineer

Frederick, MD · On-site

$56.75 - $75.25/hr

With experts in biomedical science, software engineering, and program management, we focus on ... Transportation Reimbursement Account (TRN) The Site Reliability Engineer role centers on ...

Design, deploy, and manage AWS infrastructure, including EC2, VPCs, networking, security controls ... Perform Site Reliability Engineering (SRE) functions, including automation of operational tasks ...

Design, deploy, and manage AWS infrastructure, including EC2, VPCs, networking, security controls ... Perform Site Reliability Engineering (SRE) functions, including automation of operational tasks ...

Design, deploy, and manage AWS infrastructure, including EC2, VPCs, networking, security controls ... Perform Site Reliability Engineering (SRE) functions, including automation of operational tasks ...

next page

Showing results 1-20

Site Reliability Engineer Manager information

See Maryland salary details

$10

$61

$89

How much do site reliability engineer manager jobs pay per hour?

As of Sep 8, 2026, the average hourly pay for site reliability engineer manager in Maryland is $61.86, according to ZipRecruiter salary data. Most workers in this role earn between $53.17 and $70.67 per hour, depending on experience, location, and employer.

What is a site reliability engineer manager?

A Site Reliability Engineer (SRE) Manager oversees a team of site reliability engineers tasked with maintaining the reliability, scalability, and performance of software systems. Their role combines leadership and technical expertise, focusing on automating operations, managing incidents, and ensuring high availability of services. They work closely with engineering and operations teams to implement best practices in monitoring, incident response, and system design. SRE Managers also mentor their teams, set reliability goals, and help drive a culture of continuous improvement within the organization.

What are the key skills and qualifications needed to thrive as a site reliability engineer manager?

To thrive as a Site Reliability Engineer Manager, you need expertise in systems engineering, incident management, and a strong background in software development or computer science, often supported by a bachelor’s degree or equivalent experience. Familiarity with cloud platforms (like AWS, GCP, or Azure), infrastructure as code tools (such as Terraform), monitoring systems (like Prometheus), and certifications in cloud or DevOps practices are highly valued. Strong leadership, effective communication, and problem-solving abilities help you guide teams and foster collaboration across departments. These skills and qualities ensure the stability, scalability, and reliability of critical systems while enabling teams to respond effectively to complex technical challenges.

How does a site reliability engineer manager typically balance technical leadership with team management responsibilities?

A Site Reliability Engineer Manager often splits their time between overseeing technical projects, such as system reliability improvements and incident response strategies, and managing the growth and well-being of their engineering team. This includes mentoring SREs, facilitating communication between teams, setting priorities, and ensuring that operational goals align with business objectives. Balancing these responsibilities requires strong organizational skills and a proactive approach to both technical challenges and people management. Successful managers regularly engage in hands-on problem-solving while also fostering a collaborative team environment.

What is the difference between Site Reliability Engineer Manager vs Site Reliability Engineer?

AspectSite Reliability Engineer (SRE)Site Reliability Engineer Manager
ResponsibilitiesFocuses on designing, implementing, and maintaining reliable systems and automationOversees SRE teams, manages projects, and aligns reliability goals with business objectives
Required SkillsStrong coding, system design, and troubleshooting skillsLeadership, team management, strategic planning
CertificationsGoogle Cloud, AWS certifications, Linux, scriptingSame as SRE, plus management certifications (e.g., PMP) often preferred
Work EnvironmentTechnical, hands-on with systems and automationManagerial, coordinating teams and projects

The main difference is that a Site Reliability Engineer focuses on technical system reliability, while a Site Reliability Engineer Manager oversees teams and strategic initiatives to ensure reliability goals are met across projects.

How much do site reliability engineer managers get paid?

Site Reliability Engineer Managers typically earn between $120,000 and $180,000 annually, depending on experience, location, and company size. They often oversee teams responsible for system reliability, incident response, and infrastructure automation, requiring strong leadership and technical skills.

Is a Site Reliability Engineer Manager a stressful job?

A Site Reliability Engineer Manager role can be stressful due to the responsibility of maintaining system uptime, managing incident responses, and ensuring reliability across complex infrastructure. The job often involves working under pressure, handling outages, and coordinating teams, but it also offers opportunities for problem-solving and leadership. Stress levels vary depending on company size, team structure, and workload management skills.

What are the most commonly searched types of Site Reliability Engineer jobs in Maryland?

The most popular types of Site Reliability Engineer jobs in Maryland are:

What cities in Maryland are hiring for Site Reliability Engineer Manager jobs?

Cities in Maryland with the most Site Reliability Engineer Manager job openings:

Infographic showing various Site Reliability Engineer Manager job openings in Maryland as of August 2026, with employment types broken down into 88% Full Time, 11% Part Time, and 1% Contract. Highlights an 79% Physical, 3% Hybrid, and 18% Remote job distribution, with an average salary of $128,677 per year, or $61.9 per hour.

Site Reliability Engineer

Axle

Frederick, MD • Hybrid

$56.75 - $75.25/hr

Full-time

Medical, Dental, Vision, Retirement, PTO

Re-posted 9 days ago


Job description

(ID: 2025-1135)

Axle is a bioscience and information technology company that offers advancements in translational research, biomedical informatics, and data science applications to research centers and healthcare organizations nationally and abroad. With experts in biomedical science, software engineering, and program management, we focus on developing and applying research tools and techniques to empower decision-making and accelerate research discoveries. We work with some of the top research organizations and facilities in the country including multiple institutes at the National Institutes of Health (NIH).

Benefits We Offer:

  • 100% Medical, Dental & Vision Coverage for Employees
  • Paid Time Off and Paid Holidays
  • 401K match up to 5%
  • Educational Benefits for Career Growth
  • Employee Referral Bonus
  • Flexible Spending Accounts:
    • Healthcare (FSA)
    • Parking Reimbursement Account (PRK)
    • Dependent Care Assistant Program (DCAP)
    • Transportation Reimbursement Account (TRN)

The Site Reliability Engineer role centers on modernizing and consolidating a complex multi-cloud environment across AWS, Azure, and GCP, building a scalable, secure, and observable platform from the ground up using Kubernetes, AI/ML infrastructure, and zero-trust principles. You'll combine DevOps and SRE practices to support mission-driven scientific and clinical programs, emphasizing automation, reliability, compliance, and proactive monitoring while enabling innovation through AI-driven tooling. The team culture is highly collaborative and growth-oriented, valuing experimentation, continuous learning, and cross-functional leadership, with opportunities to shape future multi-cloud and platform engineering solutions.

Responsibilities:

  • Design and implement enterprise-grade monitoring and observability frameworks (metrics, logs, traces) across distributed systems using enterprise Splunk, Grafana and Open-telemetry tools

  • Establish and manage SLIs, SLOs, and error budgets to drive reliability improvements 

  • Develop and maintain real-time asset inventory systems across cloud, on-prem, and hybrid environments 

  • Automate workload onboarding and offboarding processes, ensuring standardization and governance 

  • Track system ownership, dependencies, and lifecycle states for operational transparency

  • Build proactive detection mechanisms using AIOps and intelligent alerting to minimize incident impact

  • Design and operate scalable, resilient, and secure infrastructure platforms across cloud and hybrid environments 

  • Implement automated compliance tracking and enforcement aligned with organizational and regulatory standards (e.g., NIST, FISMA, FedRAMP) 

  • Embed ITIL processes (incident, change, problem, configuration management) into SRE workflows 

  • Build and maintain automated deployment environments and pipelines that enforce security, compliance, and operational standards 

  • Develop "golden paths" and standardized platform templates for consistent workload deployment 

  • Automate provisioning, patching, configuration management, and environment lifecycle 

  • Leverage AI/ML coding assistants and vibe coding practices to rapidly develop automation scripts, tools, and internal platforms 

  • Integrate AI-driven tooling into DevOps pipelines for code quality, security scanning, and operational insights 

  • Lead adoption of AI-enhanced SRE practices, including intelligent remediation and predictive operations

  • Champion DevOps and SRE practices including Infrastructure as Code, CI/CD, observability, and reliability engineering 

  • Build developer-friendly platforms ("golden paths") that simplify deployments, reduce friction, and improve velocity 

  • Enable and optimize infrastructure for AI/ML workloads, including data pipelines, storage systems, and inference environments, GPU-enabled and high-performance compute workloads 

  • Build and manage containerized and orchestrated platforms (Docker, Kubernetes) 

  • Support cloud migration, modernization, and platform standardization initiatives 

  • Ensure systems meet security, compliance, backup, and disaster recovery requirements 

  • Evangelize and promote best practices in DevOps, SRE, and platform engineering to developer communities

  • Stay abreast of new technologies in your areas but not limited to AIOps, MLOps, cloud computing & deployment, site reliability engineering, infrastructure automation, security best practices, data engineering etc. 

Requirements:

  • Must have total of 6+ experience DevOps / SRE roles with monitoring and observability tools (Prometheus, Grafana, ELK, or cloud-native equivalents) for on-prem and cloud hosted workloads.

  • Must have 4+ years of Hands-on Linux experience that includes Ubuntu/CentOS/Red Hat operating systems, containers, dependency management and administration support

  • Must have 4+ years of experience automating Infrastructure-as-Code (IaC) deployments to one of the following cloud platforms Amazon AWS, Google GCP and Microsoft Azure

  • Must have 4+ years with CI/CD and automation tools such as Terraform, Ansible, Chef, Puppet, Jenkins, GitHub Actions

  • Strong scripting skills (Python, Bash, PowerShell or similar)

  • Must be proficient using vibe coding and coding assistants to develop scripts, tools and applications for the DevOps and SRE use cases

  • Must have proficiency to debug or troubleshoot and/or deploying SQL and/or NoSQL databases, object storage, web servers, open-source programming stack for Node.JS, R, Python, .NET Core, Java is desired but not mandatory

  • Must be willing to learn new technologies, adopt and adapt to emerging technologies or needs from a project to a project

  • Cloud certifications is preferred 

  • Certifications in Grafana, Splunk, Docker, Kubernetes is preferred but optional

Disclaimer: The above description is meant to illustrate the general nature of work and level of effort being performed by individuals assigned to this position or job description. This is not restricted as a complete list of all skills, responsibilities, duties, and/or assignments required. Individuals may be required to perform duties outside of their position, job description or responsibilities as needed.

The diversity of Axle's employees is a tremendous asset. We are firmly committed to providing equal opportunity in all aspects of employment and will not tolerate any illegal discrimination or harassment based on age, race, gender, religion, national origin, disability, marital status, covered veteran status, sexual orientation, status with respect to public assistance, and other characteristics protected under state, federal, or local law and to deter those who aid, abet, or induce discrimination or coerce others to discriminate.

Accessibility: If you need an accommodation as part of the employment process please contact: careers@axleinfo.com

This role has a market-competitive salary with an anticipated base compensation range listed below. Actual salaries will vary depending on a candidate's experience, qualifications, skills, and location.