1

Reliability Engineer Manager Jobs in Sterling, VA

Site Reliability Engineer

Sterling, VA ยท On-site

$56.50 - $75/hr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

The SRE executes and analyzes manual IT operations/admin tasks (log analysis, performance tuning, patch management, testing, and incident response) and converts them to automated tasks. The SRE works ...

Senior Reliability Engineer

Sterling, VA

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Apply derating and thermal management analysis to high-power EW RF components; review junction ... Must have 8+ years of reliability engineering experience, including significant work on RF ...

SRE Engineer

Arlington, VA ยท On-site

$65.75 - $87.25/hr

... manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for ... Site reliability engineering, monitoring, automation, incident response, performance optimization ...

Senior Reliability Engineer

Sterling, VA ยท On-site

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Apply derating and thermal management analysis to high-power EW RF components; review junction ... Must have 8+ years of reliability engineering experience, including significant work on RF ...

Senior Reliability Engineer

Sterling, VA ยท On-site

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Apply derating and thermal management analysis to high-power EW RF components; review junction ... Must have 8+ years of reliability engineering experience, including significant work on RF ...

SRE Engineer

Arlington, VA ยท On-site

$65.75 - $87.25/hr

... manage Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets for ... Site reliability engineering, monitoring, automation, incident response, performance optimization ...

You will help ensure MC&FP systems are reliable, scalable, resilient, and efficiently managed ... Expert knowledge of site reliability engineering practices, system monitoring, incident management ...

Reliability Engineer

Ashburn, VA ยท On-site

$120 - $160/hr

The ideal candidate will lead monitoring solutions, manage ITIL engineers, automate processes, and collaborate across IT and business teams to improve service reliability. Expertise in AWS ...

Reliability Engineer, Mechanical, NA

Ashburn, VA ยท Hybrid

$90K - $120K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Reliability Engineering Department The Reliability Engineering Team is responsible for the overall ... Risk Management and Analysis: * Develop risk management plans that will anticipate reliability ...

Site Reliability Engineer - Hybrid

Reston, VA ยท On-site

$59.25 - $78.75/hr

Second round would be an in-person interview Manager's call notes * This is an SRE role. SRE is under a shared services team within Fannie Mae who works with different application teams. So, multi ...

Reliability Engineer

Washington, DC ยท On-site

$116K - $146K/yr

The ideal candidate will lead monitoring solutions, manage ITIL engineers, automate processes, and collaborate across IT and business teams to improve service reliability. Expertise in AWS ...

Site Reliability Engineer, Lead

Chantilly, VA ยท On-site

$99 - $225/hr

  • Medical

  • Life

  • Retirement

  • PTO

As a Lead Site Reliability Engineer (SRE) on our team, you'll be responsible for ensuring the ... Experience with deploying and managing OpenTelemetry . Experience with AWS CloudWatch, AWS EKS, and ...

Site Reliability Engineer IV

Sterling, VA ยท On-site

$120 - $160/hr

Strong understanding of JVM fundamentals (heap/memory management, garbage collection, OOM issues, thread analysis) * Proven experience with SRE practices, including: * Incident response and on-call ...

Reliability Engineer

Mclean, VA ยท On-site

$103K - $130K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

The Reliability Engineer will act generally as a member of a design, analysis or review team on ... Coordinate and work closely with other engineering, logistics, financial, and program management ...

Showing results 21-40

Reliability Engineer Manager information

See Sterling, VA salary details

$60.4K

$116.8K

$139.6K

How much do reliability engineer manager jobs pay per year?

As of Aug 16, 2026, the average yearly pay for reliability engineer manager in Sterling, VA is $116,834.00, according to ZipRecruiter salary data. Most workers in this role earn between $101,500.00 and $127,800.00 per year, depending on experience, location, and employer.

What does a reliability engineer manager do?

A Reliability Engineer Manager oversees teams responsible for improving the reliability and performance of systems, machinery, or processes within an organization. They develop maintenance strategies, lead root cause analyses of failures, and implement best practices to minimize downtime and costs. Additionally, they collaborate with other departments to ensure that reliability goals align with business objectives and compliance standards. Their role is crucial in industries such as manufacturing, energy, and technology, where system uptime and safety are critical.

What are some common challenges reliability engineer managers face when balancing long-term reliability improvements with immediate operational demands?

Reliability Engineer Managers often need to prioritize urgent maintenance issues while also driving long-term reliability initiatives. Balancing these competing demands can be challenging, as immediate equipment failures may require quick fixes that temporarily interrupt ongoing improvement projects. Effective managers work closely with operations, maintenance, and engineering teams to communicate priorities, allocate resources, and implement sustainable solutions that address root causes rather than just symptoms. This role typically involves using data-driven decision-making and fostering a culture of proactive maintenance and continuous improvement.

What are the key skills and qualifications needed to thrive as a reliability engineer manager?

To thrive as a Reliability Engineer Manager, you need a strong background in engineering principles, reliability analysis, and maintenance strategies, typically supported by a degree in engineering and experience in reliability roles. Familiarity with reliability-centered maintenance (RCM), failure mode and effects analysis (FMEA), and asset management software such as SAP or Maximo is common, along with certifications like Certified Reliability Engineer (CRE). Leadership, problem-solving, and effective communication are vital soft skills for managing teams and driving cross-functional initiatives. These competencies are crucial for minimizing downtime, optimizing equipment performance, and ensuring long-term operational efficiency.

What is the difference between Reliability Engineer Manager vs Reliability Engineer?

AspectReliability EngineerReliability Engineer Manager
Required CredentialsBachelor's in Engineering or related field; certifications like CRC, CRESame as Reliability Engineer, plus leadership experience
Work EnvironmentDesign, analyze, and improve system reliability; often in teamsOversees Reliability Engineers; manages projects and teams
Employer & Industry UsageManufacturing, aerospace, energy, automotiveSame industries, with added managerial responsibilities
Common Search & ComparisonFocuses on technical skills and hands-on reliability tasksFocuses on leadership, team management, and strategic planning

The main difference between a Reliability Engineer and a Reliability Engineer Manager lies in their responsibilities. The Reliability Engineer focuses on technical analysis and system improvements, while the Reliability Engineer Manager oversees teams, manages projects, and develops strategies to enhance reliability across the organization.

What cities near Sterling, VA are hiring for Reliability Engineer Manager jobs?

Cities near Sterling, VA with the most Reliability Engineer Manager job openings:

Site Reliability Engineer

Nwis

Sterling, VA โ€ข On-site

$56.50 - $75/hr

Full-time

Medical, Dental, Vision, Retirement, PTO

Re-posted 25 days ago


Job description

Nightwing provides technically advanced full-spectrum cyber, data operations, systems integration and intelligence mission support services to meet our customers' most demanding challenges. Our capabilities include cyber space operations, cyber defense and resiliency, vulnerability research, ubiquitous technical surveillance, data intelligence, lifecycle mission enablement, and software modernization. Nightwing brings disruptive technologies, agility, and competitive offerings to customers in the intelligence community, defense, civil, and commercial markets.


Job Title: Site Reliability Engineer
Location:Sterling, VA
Clearance:TS/SCI Poly

**This position is CONTINGENT upon contract award**

The Site Reliability Engineer (SRE) collaboratively works closely with the contract leadership, Platform teams, and Sponsor to refine the operational and technical strategy to automate key portions of IT operations and enable the Product team (Platform) to bring new software or new features to production as quickly as possible. The SRE executes and analyzes manual IT operations/admin tasks (log analysis, performance tuning, patch management, testing, and incident response) and converts them to automated tasks. The SRE works with the Platform, Network and Data Operations teams to assist in deployment planning and onboard systems. They assist with monitoring, system analysis, and IT operations support. Daily tasks include, but are not limited to:

  • Work with Sponsor, Mission partners, and technical personnel to deliver robust scalable operations architecture that meets the customer goals for the enterprise.
  • Analyze, define, and document requirements for data, workflow, logical processes, hardware and operating system environment, and network connectivity, other system interfaces, internal and external checks and controls, and outputs.
  • Monitor and track metrics, logs and traces across all services in the system/network and provide context for identifying root causes in the event of an incident, performance degradation, or availability issue.
  • Perform Network/Cloud optimization and resilience planning
  • Develop capabilities to automate hardware/software provisioning, monitoring, patching, and troubleshooting.
  • Collaborate with and assist Platform team and leadership in network and security health, intrusions or inappropriate activities.
  • Optimize business processes, workflows, and service operations by building efficient on-call processes and streamlining alerting workflows.
  • Leverage operational data to automate systems administration, operations and incident response processes to improve enterprise reliability to manage IT environment complexity.
  • Works with LSA, Lab Manager, and CM to compose technical documents including Design, Deployment, System specifications and Host Nation baselines, updates, user's manuals, training materials, installation guides, proposals, and reports.
  • Work with the OM to implement ITSM best practices for ICA/Service discrepancy and reporting, issue resolution and operations support to include Tier 2/3 escalation.

Required Skills:

  • Programming: Proficiency in at least one programming language (e.g., Python, Go, Java, or JavaScript) is essential for automating tasks and developing tools.
  • Linux/Unix Systems Administration: Strong knowledge of Linux/Unix operating systems, including command-line tools and system administration tasks.
  • Networking: Understanding of network protocols, infrastructure, and troubleshooting techniques.
  • Database Management: Familiarity with database technologies and principles.
  • Automation: Experience with automation tools and techniques, such as configuration management (e.g., Ansible, Puppet, Chef) and orchestration (e.g., Kubernetes).
  • Monitoring and Logging: Experience with monitoring tools and logging systems.
  • Problem-Solving: Strong analytical and problem-solving skills to diagnose and resolve system issues.
  • Communication: Ability to communicate technical information clearly and concisely to both technical and non-technical audiences.
  • Collaboration: Ability to work effectively with cross-functional teams, including software developers and operations personnel.

Desired Skills:

  • Cloud Technologies: Experience with cloud platforms (e.g., AWS, Google Cloud, Azure).
  • Containerization: Knowledge of containerization technologies (e.g., Docker, Kubernetes).
  • DevOps Principles: Understanding DevOps principles and practices.
  • Service Level Objectives (SLOs) and Service Level Agreements (SLAs): Experience with defining, tracking, and managing SLOs and SLAs.
  • Data Analysis: Experience with data analysis and visualization tools.

Desired Certs:

  • Global Skill Development Council (GSDC) Site Reliability Engineering (SRE) Foundation Certification (CSREF).
  • AWS Certified SysOps Administrator - Associate.
  • Google Cloud Certified Professional Cloud Architect.
  • Azure Certified Solutions Architect Expert.


Salary & Benefits

The salary associated with this position ($122,000-$253,000) is commensurate with the selected candidate's qualifications, years of relevant experience, and demonstrated level of expertise. Compensation will be determined based on these factors to ensure alignment with skills, responsibilities, and market standards.


Nightwingoffers medical, vision and dental insurance coverage in addition to a 401k plan, PTO, Holidays, and additional insurances.


Nightwing is An Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability or veteran status, age or any other federally protected class.