1

Director Observability Jobs in Virginia (NOW HIRING)

AIOps Engineer

Fort Belvoir, VA · On-site

$190K - $218K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Experience leading technical working groups and directing the efforts of adjacent infrastructure ...

Senior Systems Integration Engineer

Reston, VA

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Engineer for resilience, observability, auditability, latency, and intermittent/low-bandwidth ... Ability to work effectively in a highly collaborative engineering environment with direct ownership ...

Senior AI OPS Engineer

Fort Belvoir, VA · On-site

$150K - $174K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Experience leading technical working groups and directing the efforts of adjacent infrastructure ...

API Gateway Engineer (Kong)

Herndon, VA · On-site +1

$170K - $210K/yr

... validation, observability integration, and knowledge transfer. Key Responsibilities: * Provide ... Direct Kong Enterprise federal implementation experience The salary compensation range for this ...

Senior AI OPS Engineer

Fort Belvoir, VA · On-site

$150K - $174K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Experience leading technical working groups and directing the efforts of adjacent infrastructure ...

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Experience leading technical working groups and directing the efforts of adjacent infrastructure ...

Quality, Security & Observability · Embed data quality rules, unit/integration tests, and ... avoid direct system scraping. · Produce "readme" docs, data dictionaries, runbooks, and post ...

Showing results 41-60

Director Observability information

What is the difference between Director Observability vs Site Reliability Engineer?

AspectDirector ObservabilitySite Reliability Engineer
Primary FocusOversees observability strategies, tools, and teams to ensure system visibility and performanceBuilds and maintains reliable systems, automates deployment, and manages incident response
CredentialsTypically requires advanced knowledge of monitoring, cloud platforms, and leadership experienceOften has software engineering background, with skills in scripting, automation, and systems engineering
Work EnvironmentLeads teams in tech companies, focusing on monitoring and analytics toolsWorks closely with development and operations teams to ensure system reliability

While both roles focus on system performance and reliability, the Director Observability primarily manages observability strategies and teams, whereas the Site Reliability Engineer is hands-on, building and maintaining reliable systems. The roles complement each other in ensuring optimal system performance and uptime.

What are the key skills and qualifications needed to thrive as a director of observability, and why are they important?

To thrive as a Director of Observability, you need deep expertise in monitoring, logging, and distributed systems, typically backed by a degree in computer science or a related field and extensive experience in IT or DevOps leadership roles. Proficiency with observability tools such as Prometheus, Grafana, Datadog, Splunk, and APM solutions, along with knowledge of cloud platforms and relevant certifications, is essential. Strong leadership, strategic thinking, and communication skills help drive cross-functional initiatives and foster a culture of reliability. These skills and qualities are crucial for ensuring system health, rapid incident response, and alignment between technical teams and organizational objectives.

How does a director of observability typically collaborate with engineering and operations teams to drive organizational goals?

A Director of Observability works closely with engineering and operations teams to ensure that systems are monitored effectively and issues are identified and resolved quickly. This collaboration often involves developing unified monitoring strategies, aligning observability tools and processes, and facilitating incident response post-mortems. The Director also leads cross-functional meetings to establish best practices, set key performance indicators (KPIs), and ensure observability is integrated into the software development lifecycle. By acting as a bridge between technical teams, they help foster a culture of transparency, reliability, and continuous improvement.

What does a director of observability do?

A Director of Observability leads the strategy and implementation of monitoring, logging, and tracing systems to ensure the health and performance of technical infrastructure. They work with engineering and operations teams to develop best practices, select appropriate tools, and set standards for observability across the organization. Their goal is to provide visibility into system behavior, quickly identify and resolve incidents, and support continuous improvement in system reliability and performance.

What are the most commonly searched types of Observability jobs in Virginia?

The most popular types of Observability jobs in Virginia are:

What are popular job titles related to Director Observability jobs in Virginia?

For Director Observability jobs in Virginia, the most frequently searched job titles are:

What job categories do people searching Director Observability jobs in Virginia look for?

The top searched job categories for Director Observability jobs in Virginia are:

What cities in Virginia are hiring for Director Observability jobs?

Cities in Virginia with the most Director Observability job openings:

AIOps Engineer

Career Listings

Fort Belvoir, VA • On-site

$190K - $218K/yr

Full-time

Medical, Dental, Vision, Retirement, PTO

Re-posted 12 days ago


Job description

Benefits:
  • 401(k)
  • 401(k) matching
  • Dental insurance
  • Health insurance
  • Paid time off
  • Profit sharing
  • Training & development
  • Tuition assistance
  • Vision insurance

Primary Responsibilities:

  • Cross-Functional Leadership: Lead the AIOps platform initiative by acting as the primary technical liaison to existing Network Engineering, ServiceNow, and SolarWinds administration teams to establish unified telemetry pipelines.
  • ITSM Orchestration & Automation: Architect closed-loop remediation workflows by deeply integrating Splunk ITSI alerts with ServiceNow Event Management and Incident Management modules.
  • Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified and classified enclaves to provide real-time situational awareness.
  • Infrastructure Telemetry Integration: Normalize and correlate network performance and fault data from SolarWinds with server and application logs to provide a holistic view of enterprise health.
  • Advanced ML Development: Deploy custom machine learning models via Splunk MLTK to identify anomalous behavior, potential cyber threats, and infrastructure degradations.
  • Secure Data Integration: Engineer secure data ingestion pipelines for telemetry data from cross-domain solutions and tactical edge devices.
  • Incident Reduction: Utilize IT Service Intelligence (ITSI) to correlate multi-source events, reducing noise and prioritizing high-impact mission alerts.
  • Cyber Defense Support: Collaborate with the Cyber Security Service Provider (CSSP) to integrate AIOps insights into defensive cyber operations (DCO).
  • Compliance & Documentation: Ensure all observability tools comply with DoW STIGs and IL5/IL6 protocols; develop and maintain architectural documentation and compliance traceability.
  • Mission Alignment: Stay current on AIOps and related capabilities relevant to DoD, federal, and intelligence mission systems.
Required Qualifications:

  • Security Clearance: Active Top Secret / Sensitive Compartmented Information (TS/SCI) required at time of hire.
  • Certification: Active IAT Level II certification (e.g., Security+ CE, CySA+, GSEC, or SSCP) required.
  • Citizenship: United States Citizenship is required.
  • Platform Experience: 7+ years of experience with Splunk Enterprise, including architectural design, cluster management, and advanced Search Processing Language (SPL).
  • AIOps & ITSM: 3+ years of experience implementing AIOps workflows, including integration with enterprise ITSM solutions (ServiceNow) for automated root cause analysis and remediation.
  • Machine Learning: Proven track record of building, testing, and tuning supervised and unsupervised models within the Splunk MLTK.
  • Scripting & Automation: Advanced scripting skills for developing custom search commands, API integrations, and automating remediation tasks (e.g., Python).
  • Leadership: Experience leading technical working groups and directing the efforts of adjacent infrastructure and development teams.
  • Operational Experience: Prior experience working within a DoW/DoD Operations Center (NOC/SOC) or supporting mission-critical systems and networks.
  • Communication: Must be able to present designs, plans, and analyses of alternatives to technical leadership boards for approvals.
Desired Qualifications:
  • Enterprise Aggregation: Experience aggregating and correlating telemetry from diverse tools, specifically SolarWinds, ServiceNow, and VMware vCenter.
  • Expert Certification: Splunk Enterprise Certified Architect or Splunk ITSI Certified Admin.
  • Cloud Observability: Experience with Cloud Native Computing Foundation (CNCF) observability tools in secure hybrid multi-cloud environments (Azure/AWS).
  • RMF/ATO Knowledge: Understanding of the Risk Management Framework (RMF) and the Authorization to Operate (ATO) process for AI/ML workloads.