2

Remote Observability Engineer Jobs in Virginia (NOW HIRING)

Fullstack Engineer

Richmond, VA · On-site +1

$115K - $141K/yr

Implement monitoring, alerting, and observability best practices * Support operational excellence ... Remote opportunities are available to candidates throughout the United States. Salary Range: $115 ...

Lead Data Engineer 3624301

Richmond, VA · Remote

$114K - $187K/yr

What's In Store For You This is a full-time, remote opportunity for candidates residing in approved ... Experience implementing automated testing, validation, monitoring, alerting, observability, and SLA ...

Senior DevOps Engineer

Mclean, VA · Remote

$131K - $168K/yr

Remote with Preference for Greenbelt, MD Job number : 854 Travel: Occasional Travel Required for ... reliability, observability, monitoring, and operational efficiency • Develop and maintain ...

... s Full-Stack Engineer with expertise in IaC (Terraform), Helm, MySQL, Kubernetes, and CI/CD ... Familiarity with observability tooling (Prometheus/Grafana, ELK Stack or similar) * Knowledge of ...

Senior DevOps Engineer

Mclean, VA · Remote

$131K - $168K/yr

Senior DevOps Engineer Job number: 842 This is a remote position. Ad Hoc is a technology company ... Familiarity with JavaScript, Node.js, Postgres, and observability tooling To learn more about ...

next page

Showing results 1-20

Remote Observability Engineer information

What are the typical collaboration patterns for a Remote Observability Engineer working with distributed teams?

Remote Observability Engineers frequently collaborate with software developers, DevOps teams, and IT operations to ensure systems are monitored effectively and issues are detected early. Working remotely, you'll often use communication tools like Slack, Jira, and video conferencing to coordinate incident response, discuss monitoring strategies, and review system health dashboards. Regular sync meetings and asynchronous updates are common, and you'll likely contribute to documentation and knowledge sharing to keep all stakeholders informed. Building strong communication habits is important, as much of the troubleshooting and improvement work hinges on clear coordination with multiple teams.

What are the key skills and qualifications needed to thrive as a Remote Observability Engineer, and why are they important?

To thrive as a Remote Observability Engineer, you need strong expertise in monitoring, logging, and tracing systems, along with a background in computer science or related technical fields. Familiarity with tools like Prometheus, Grafana, ELK Stack, Datadog, and cloud platforms is typically required, as well as relevant certifications such as AWS Certified Cloud Practitioner or Google Cloud Professional DevOps Engineer. Excellent problem-solving abilities, communication skills, and a proactive mindset help you detect and resolve issues before they impact users. These competencies ensure system reliability, enable rapid incident response, and support seamless collaboration in distributed environments.

What is the difference between Remote Observability Engineer vs Site Reliability Engineer?

AspectRemote Observability EngineerSite Reliability Engineer
CredentialsKnowledge of monitoring tools, scripting, cloud platformsSame as Observability Engineer, plus SRE certifications often preferred
Work EnvironmentFocus on monitoring, logging, and tracing systems remotelyBroader scope including system reliability, incident response, and automation
Industry UsagePrimarily in tech, SaaS, cloud servicesWidely in tech, finance, and large-scale online services

The Remote Observability Engineer specializes in monitoring and analyzing system performance remotely, focusing on tools like logs and metrics. In contrast, the Site Reliability Engineer has a broader role, ensuring overall system reliability, automation, and incident management. While both roles require similar technical skills, SREs often have additional responsibilities related to system resilience and scalability.

What is a Remote Observability Engineer?

A Remote Observability Engineer is a professional responsible for designing, implementing, and maintaining systems that monitor the health, performance, and reliability of software applications and infrastructure from a remote location. They use observability tools to collect and analyze logs, metrics, and traces, helping organizations quickly detect and resolve issues. Their work ensures that distributed systems are transparent, reliable, and efficient, often collaborating with development, operations, and security teams. Remote Observability Engineers often work from anywhere, leveraging cloud-based tools and platforms to manage complex IT environments.
What are the most commonly searched types of Observability Engineer jobs in Virginia? The most popular types of Observability Engineer jobs in Virginia are:
What job categories do people searching Remote Observability Engineer jobs in Virginia look for? The top searched job categories for Remote Observability Engineer jobs in Virginia are:
What cities in Virginia are hiring for Remote Observability Engineer jobs? Cities in Virginia with the most Remote Observability Engineer job openings:
Infographic showing various Remote Observability Engineer job openings in Virginia as of July 2026, with employment types broken down into 80% Full Time, and 20% Contract. Highlights an 100% Remote job distribution.
Cloud Infrastructure Engineer - Kubernetes (Remote)

Cloud Infrastructure Engineer - Kubernetes (Remote)

Oxley Enterprises, Inc.

Stafford, VA • Remote

$107K - $140K/yr

Full-time

Medical, Dental, Vision, Life, Retirement

Posted 26 days ago


Job description

The following states/districts are excluded from this job ad: AK, CA, CO, CT, DC, HI, LA, MA, MN, MO, NE, NV, NH, NJ, NM, NY, ND, OR, PR, RI, VT, WA, WY

Future Need - Actively Interviewing

Location: Remote in any United States jurisdiction not excluded from this job advertisement.

Keep the engine running on a complex Kubernetes environment! As a Kubernetes Platform Engineer, you will operate and maintain AWS EKS clusters supporting 300+ applications across production, staging, and sandbox environments.

Position Description: The Cloud Infrastructure Engineer - Kubernetes supports the day-to-day operation, maintenance, and continuous improvement of AWS EKS clusters including cluster lifecycle, node operations, add-on version compliance, namespace administration, Istio service mesh operations, and Infrastructure as Code (IaC)-driven configuration.

Minimum/General Experience: 3 years of experience in Kubernetes platform engineering and managing AWS EKS clusters in production environments

Minimum Education: Bachelor's Degree in computer science, information technology, systems engineering, or related field

Essential Skills/Qualifications:

  • Excellent experience managing AWS EKS clusters including managed node groups, Fargate profiles, cluster upgrades, add-on lifecycle management, and multi-cluster operations at enterprise scale
  • Excellent knowledge of Kubernetes cluster operations including namespace administration, resource quotas, limit ranges, pod disruption budgets, HPA/VPA, and cluster autoscaler configuration
  • Excellent ability to detect and remediate node failures within 30 minutes using cordon, drain, and replace procedures
  • Excellent experience with Istio service mesh operations including control plane management, mTLS enforcement, traffic management, virtual services, and sidecar injection policy on EKS
  • Excellent experience implementing and maintaining 100% IaC governance using Terraform for all cluster configuration
  • Excellent knowledge of Kubernetes add-on lifecycle including CoreDNS, kube-proxy, CNI, AWS Load Balancer Controller, and EBS/EFS CSI drivers
  • Above average experience with Kubernetes monitoring and observability including Prometheus, Grafana, Dynatrace, log aggregation, and distributed tracing
  • Working knowledge of EKS multi-region deployment patterns including cluster federation and cross-region service discovery
  • Experience supporting a federal agency
  • Excellent verbal and written communication skills

General Physical Requirements needed to perform the essential functions of this job may vary based on the location of the assignment.

  • Assignment Location - Remote
  • Sedentary Work - Exerting up to 10 pounds of force occasionally and/or a negligible amount of force frequently or constantly to lift, carry, push, pull or otherwise move objects.
  • Typing, communicating, repetitive motions.
  • Close visual acuity to prepare and analyze data, view computer monitors and read. May need to view presentation screens and other visual aids in a virtual setting.
  • Inside environmental conditions with protection from outside elements.

Security: Active Federal Civilian Public Trust clearance

  • U.S. Citizenship or Permanent Resident that has lived in the United States for at least 3 years

Federal Civilian Public Trust Consists of a review of up to but not limited to:

  • Covers 10 year period and in some instances lifetime events
  • OPM Security Investigations Index (SII)
  • DOD Defense Central Investigations Index (DCII)
  • National Agency Check (NAC) records
  • FBI name check
  • FBI fingerprint check
  • Credit report check
  • Written inquiries to previous employers and references listed on the application for employment
  • Potential interviews with the subject, spouse, neighbors, supervisor, coworkers
  • Law enforcement check
  • Court records check
  • Education check - Attendance and Degrees

Acceptable Credentials

Tasks/activities include, but are not limited to:

  • Operates and maintains all AWS EKS clusters across production, staging, and sandbox
  • Detects and remediates node failures within 30 minutes using cordon, drain, and replace
  • Manages EKS cluster add-on lifecycle ensuring 100% of add-ons (e.g., CoreDNS, kube-proxy, CNI, AWS Load Balancer Controller, EBS/EFS CSI) remain within one minor version of the EKS cluster version with no EOL/EOS components in production
  • Implements and maintains 100% IaC governance using Terraform for all cluster configuration and namespace management
  • Prohibits manual production changes except through approved break-glass procedures
  • Performs weekly drift detection
  • Manages Istio service mesh operations including control plane upgrades, mTLS policy enforcement, traffic routing, and sidecar injection governance
  • Configures and maintains autoscaling policies ensuring no production workload experiences resource saturation exceeding 80% for more than 5 minutes
  • Participates in daily Change Control Board (CCB) meetings for all Kubernetes cluster and namespace changes
  • Conducts post-implementation validation within 2 hours of each production change
  • Ensures all platform components are integrated with centralized logging and monitoring
  • Ensures no platform change is implemented without sufficient post-deployment production testing
  • Supports the Scheduling Event Bus delivery by providing Kubernetes infrastructure design, namespace provisioning, and operational support for event-driven workloads
  • Contributes EKS cluster health metrics, add-on compliance status, node utilization data, and incident reports to the Monthly Maintenance Report

Compensation & Benefits: The annual projected pay range for this position is $66,923 - $110,863 with consideration being given to various factors including but not limited to qualifications, experience, job responsibilities, and geographic location.

Oxley Enterprises, Inc. offers a full array of benefits including:

  • Medical, dental, vision and prescription drug coverage for you and your family.
  • Life Insurance, short-term disability and long-term disability paid for by the Company.
  • Supplemental coverages including Accident, Critical Illness, and Hospital.
  • Additional Life insurance coverage for you and your dependents.
  • 401k plan with various options to select based on your retirement goals.

Oxley Enterprises, Inc. is a certified service-disabled veteran-owned (SDVOSB), veteran-owned (VOSB), and woman-owned small business (WOSB) that has 26 years of experience building and delivering quality IT systems and programs. Oxley is ranked in the INC 5000 7 times (2016, 2017, 2018, 2021, 2023, 2024, 2025). Oxley is a 2019 - 2025 Department of Labor HIRE Vets Medallion Award Winner. Oxley is Virginia Values Veterans certified.

All qualified applicants will receive consideration for employment without regard to any status protected by applicable federal, state, or local law.

If you require a reasonable accommodation to apply for a position at Oxley Enterprises, Inc., please send an email to our Human Resources Department at: careers@oxleyenterprises.com with the following information:

Subject Line: Accommodation Request

Provide a description of your accommodation request

Include your contact information: Full name, Email address, Best number to reach you (optional)

We participate in the E-Verify program. http://www.dhs.gov/E-Verify