1

Sre Observability Engineer Jobs in Virginia (NOW HIRING)

Staff Site Reliability Engineer

Reston, VA · On-site

$59.25 - $78.75/hr

The Site Reliability Engineering team drives reliability strategy, elevates engineering standards ... Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog ...

Staff Site Reliability Engineer

Williamsburg, VA · On-site

$54.75 - $72.75/hr

About the Role We're hiring a Staff Site Reliability Engineer to define and strengthen how ... Build world-class observability: Develop a cohesive approach to metrics, logs, traces, dashboards ...

Sr Site Reliability Engineer

Springfield, VA · On-site

$59.75 - $79.50/hr

In this role, you'll join our SRE team to help keep our platforms running smoothly, improve our observability and incident response capabilities, and partner with development teams to deliver ...

Senior Observability Engineer

Mclean, VA · On-site

$107K - $147K/yr

Strong understanding of observability architecture, SRE principles, and cloud-native monitoring practices. * Hands-on experience with AWS cloud platforms. * Knowledge of anomaly detection, Open ...

Senior Site Reliability Engineer

Mclean, VA · On-site

$58.50 - $77.75/hr

As a Senior Site Reliability Engineer, you will play a key role in designing, operating, and ... Design observability strategies using monitoring, logging, tracing, and alerting platforms.

Senior Site Reliability Engineer

Mclean, VA · On-site

$57.50 - $76.50/hr

As a Senior Site Reliability Engineer, you will play a key role in designing, operating, and ... Design observability strategies using monitoring, logging, tracing, and alerting platforms.

Observability Engineer II

Mclean, VA · Hybrid

$79K - $151K/yr

Bachelor's degree or equivalent practical experience. * 6+ years of experience in observability, monitoring, SRE, or platform operations roles. * Strong hands-on experience with Dynatrace and ...

Senior Site Reliability Engineer

Mclean, VA · On-site

$58.50 - $77.75/hr

As a Senior Site Reliability Engineer, you will play a key role in designing, operating, and ... Design observability strategies using monitoring, logging, tracing, and alerting platforms.

Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability ... Design, implement, and own observability infrastructure including metrics, logging, tracing, and ...

Site Reliability Engineer - CTJ - POLY

Reston, VA · On-site

$59.25 - $78.75/hr

Builds and improves observability (metrics, logs, traces, dashboards, alerts) and uses it to detect ... Site Reliability Engineering IC4 - The typical base pay range for this role across the U.S. is USD ...

Requirements As a Senior Site Reliability Engineer at Commence, you will own the reliability ... Design, implement, and own observability infrastructure including metrics, logging, tracing, and ...

As a Senior Site Reliability Engineer at Commence, you will own the reliability, scalability, and ... Design, implement, and own observability infrastructure including metrics, logging, tracing, and ...

Showing results 41-60

Sre Observability Engineer information

What is an SRE Observability engineer?

An SRE Observability Engineer is a professional responsible for designing, implementing, and maintaining systems that provide visibility into the health, performance, and reliability of IT infrastructure and applications. They focus on monitoring, logging, and tracing to detect and resolve issues proactively. Their work ensures that systems are observable, meaning engineers can understand and diagnose problems quickly, improving uptime and user experience. SRE Observability Engineers often collaborate with development and operations teams to establish best practices for observability tools and processes.

How does an SRE Observability engineer typically collaborate with software development and IT operations teams?

As an SRE Observability Engineer, you will frequently partner with software development and IT operations teams to design, implement, and refine monitoring and alerting systems. You’ll help developers identify key application metrics, guide teams in instrumenting code, and ensure that operational data is accessible and actionable. Close collaboration is essential for resolving incidents quickly and for proactively improving system reliability, making communication and teamwork crucial skills in this role.

What are the key skills and qualifications needed to thrive as an SRE Observability engineer, and why are they important?

To thrive as an SRE Observability Engineer, you need a strong background in systems engineering, monitoring, troubleshooting, and performance analysis, often supported by a degree in computer science or related field. Familiarity with observability tools such as Prometheus, Grafana, ELK Stack, and cloud platforms like AWS or GCP, along with certifications like AWS Certified DevOps Engineer, is highly beneficial. Strong problem-solving abilities, attention to detail, and effective communication are vital soft skills in this role. These competencies are crucial for ensuring system reliability, early detection of issues, and efficient collaboration across teams.

What is the difference between Sre Observability Engineer vs Sre Reliability Engineer?

AspectSre Observability EngineerSre Reliability Engineer
Primary FocusImplementing and maintaining observability tools, monitoring, and alerting systemsEnsuring system reliability, availability, and performance
Skills & CertificationsMonitoring tools, scripting, cloud platforms, observability frameworksSystem architecture, incident management, automation, cloud expertise
Work EnvironmentDevOps teams, monitoring and operations departmentsProduction systems, incident response teams

While both roles focus on system health, the Sre Observability Engineer specializes in building and managing observability tools, whereas the Sre Reliability Engineer concentrates on overall system reliability and incident management. Both roles often collaborate but have distinct core responsibilities.

What cities in Virginia are hiring for Sre Observability Engineer jobs?

Cities in Virginia with the most Sre Observability Engineer job openings:

Staff Site Reliability Engineer

Reston, VA • On-site

Transunion
IT Services • 5 - 10K employees

$59.25 - $78.75/hr

Full-time

Medical, Dental, Vision, PTO

Re-posted 26 days ago


Key responsibilities

  • Drive reliability strategy, lead high-risk technical initiatives, and set engineering standards for the platform.

  • Participate in on-call rotation, respond to incidents, and lead high-impact maintenance events with minimal customer impact.

  • Contribute to architectural and strategic decisions, research and implement new systems and tooling, and elevate team standards through process improvements.


TransUnion rating

9.3

Company rating: 9.3 out of 10

Based on 8 frontline employees who took The Breakroom Quiz


Job description

TransUnion's Job Applicant Privacy Notice

Personal Information We Collect

Your Privacy Choices

Team Overview

At TransUnion, this role will report to a DevOps Director. The Site Reliability Engineering team drives reliability strategy, elevates engineering standards, and owns some of the most complex and consequential work on the platform.
As a Staff Site Reliability Engineer at TransUnion, you will serve as a senior technical leader and force multiplier on the SRE team. Operating with full autonomy, you will drive reliability strategy, lead high-risk technical initiatives, and set the engineering standards that elevate the entire team. You'll bring deep expertise across GCP, Kubernetes, CI/CD pipelines, and monitoring platforms - contributing to strategic decisions on major platform components while fully participating in on-call rotation. Whether stepping in to lead the team, owning complex capacity and security work, or anchoring incident response with calm and maturity, your impact will be felt across the platform and the people around you. This is a hybrid position and involves regular performance of job responsibilities virtually as well as in-person at an assigned TU office location for a minimum of two days a week.

Role Overview and Core Responsibilities

Technical Leadership & Strategic Influence

  • Recognized expert across multiple systems; actively contributes to architectural and strategic decisions around major platform components.
  • Leads research, testing, implementation, and continuous improvement for new systems and tooling.
  • Performs complex, high-impact work including capacity planning, load testing, and security improvements.

Operational Excellence & On-Call

  • Fully participates in the team's on-call rotation; models calm, effective, and blameless incident response.
  • Serves as a significant technical contributor during major incidents and problem resolution.
  • Plans and leads high-risk maintenance events with minimal to no customer impact.

Standards & Team Elevation

  • Elevates team standards through new tooling, processes, procedures, and effective communication.
  • Capable of stepping in to lead and represent the team - a trusted resource during transitions or coverage gaps.
  • Sets new professional benchmarks in technical quality, engineering culture, and cross-functional collaboration.

Required Knowledge and Experiences

  • 5+ years of experience in Cloud Architecture, Site Reliability Engineering, Platform Engineering, or related fields - with a proven track record of designing and delivering at enterprise scale.
  • Deep, hands-on expertise with Google Cloud Platform (GCP) and Kubernetes (K8s) - running high-volume, high-availability workloads with 99.999% reliability targets.
  • Mastery of CI/CD pipeline architecture - designing end-to-end delivery systems that are fast, safe, and built for scale.
  • Expert-level command of monitoring, observability, and alerting platforms (e.g., Datadog, Prometheus, Grafana, PagerDuty) - you define what good looks like.
  • Deep Linux expertise - from kernel internals and system performance tuning to hardening and troubleshooting at the OS level in production environments.
  • Strong command of database architecture - including relational (PostgreSQL, MySQL, Cloud SQL) and NoSQL (Bigtable, Firestore, Redis) systems, with experience designing for high availability, replication, failover, and performance at scale.
  • Advanced networking knowledge - including VPCs, subnets, DNS, load balancing, firewall rules, VPNs, private service connect, and hybrid connectivity patterns across cloud and on-prem environments.
  • Proven expertise in Infrastructure-as-Code (IaC) - designing and enforcing scalable, reusable frameworks using Terraform, Pulumi, or equivalent tools.
  • Strong proficiency in scripting and automation (e.g., Python, Bash, Go) - building the tools and workflows that eliminate toil and accelerate delivery.
  • Hands-on experience designing and integrating AI/ML-powered solutions into cloud-native platforms - including familiarity with LLM orchestration, vector databases, model serving infrastructure, and AI observability - with the ability to evaluate emerging tools and translate them into reliable, production-grade capabilities.

Benefits that support every part of your life:

At TransUnion, we design benefits to help youfeel well, do well, and plan well-from day one.

For Your Health: Enjoyday-one eligibilityfor medical, dental, and vision coverage, plus supplemental plan options. Spousal, domestic partner, and other eligible dependent coverage is available on select plans. Choose taxadvantagedHSAandFSAaccounts to make everyday care more affordable.

For Your Protection: We've got your back withcompanypaid basic life and AD&D, optionalvoluntary life and AD&Dfor you and your family, andshort and longterm disability. You can also opt into alegal plan,pet insurance, andtravel accident coverage.

For Your Family: Fromadoption assistance and fertility planning coveragetocaregiver support, we're here for every chapter. AccessDependent Care FSA for possibility of an employer match, a complimentaryCare@Workmembership, andup to 12 weeks of paid parental leavewith eligibility for a thoughtful, gradual return.

For Your Future: Build toward what's next with our401(k) with employer matchandEmployee Stock Purchase Plan (ESPP). Tapfinancial wellness resources,career coaching, and optionallongterm care insuranceto plan confidently.

For You: Grow and recharge withtuition reimbursement,flexible time off for exempt employees or paid time off for nonexempt employees, up to 12 paid holidays per year, commuter benefits, employeediscounts,charitable gift matching, andpaid volunteer time off, plus corporate volunteer events that make it easy to give back.

For Your Wellness: Access24/7 supportincluding professionaltherapy,coaching, and emotional wellbeing programs alongside guided meditation and resources that supportphysical, mental, social, and financial wellness.

We are committed to being a place where diversity is not only present, it is embraced. As an equal opportunity employer, all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, age, disability status, veteran status, genetic information, marital status, citizenship status, sexual orientation, gender identity or any other characteristic protected by law. Additionally, in accordance with Section 503 of the Rehabilitation Act of 1973 and the Vietnam Era Veterans' Readjustment Assistance Act of 1974, TransUnion takes affirmative action to employ and advance in employment qualified individuals with a disability and protected veterans in all levels of employment and develops annual affirmative action plans. Components of TransUnion's Affirmative Action Program for individuals with disabilities and protected veterans are available for review to any associate or applicant for employment upon request by contacting ERCoE@transunion.com.

Qualified applicants with arrest or conviction records will be considered for employment in accordance with applicable law, including the Los Angeles County Fair Chance Ordinance for Employers, the San Francisco Fair Chance Ordinance, Fair Chance Initiative for Hiring Ordinance, and the California Fair Chance Act.

Adherence to Company policies, sound judgment and trustworthiness, working safely, communicating respectfully, and safeguarding business operations, confidential and proprietary information, and the Company's reputation are also essential expectations of this position.

Pay Scale Information:The U.S. base salary range for this position is $112,500.00 - $187,500 annually. *The salary range for this position reflects a reasonable estimate of the range of compensation for this job. At TransUnion, actual compensation is based on careful consideration of additional factors such as (but not limited to) an individual's education, training, work experience, job-related skill set, location, and industry knowledge, as well as the scope and responsibilities of the position and market considerations. Regular, fulltime non-sales positions may be eligible to participate in TransUnion's annual bonus plan. Certain positions may be also eligible for long-term incentives and other payments based on applicable company guidance and plan documents.

TransUnion Overview:

At TransUnion, we encourage and are committed to creating a real, positive impact and shared sense of purpose within our Workforce for Good, which empowers our people to grow, innovate and contribute to a better future for our communities and customers. We strive to build an environment where our associates are in the driver's seat of their professional development- while having access to help along the way. We recognize that success comes when our associates thrive both professionally and personally; that's why we prioritize work/life flexibility and offer resources for our teams across the globe to collaborate and drive excellence.

Be a part of our Workforce for Good - you'll work with great people, pioneering products and cutting-edge technology.

TransUnion's Internal Job Title:

Architect, Development Ops

Company:

TransUnion LLC

What TransUnion employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom