1

Observability Aiops Engineer Jobs in Virginia (NOW HIRING)

Monitoring and Observability : Build and maintain dashboards using Grafana, Prometheus, or Kibana ... AIOps. About Us Formed through the strategic union of Sev1Tech and ERT, Entarian is a premier ...

Monitoring and Observability : Build and maintain dashboards using Grafana, Prometheus, or Kibana ... AIOps. About Us Formed through the strategic union of Sev1Tech and ERT, Entarian is a premier ...

Monitoring and Observability : Build and maintain dashboards using Grafana, Prometheus, or Kibana ... AIOps. Formed through the strategic union of Sev1Tech and ERT, Entarian is a premier provider of ...

Senior AI Engineer

Reston, VA · Remote

$127K - $168K/yr

Integrate AI into IT operations (ticket triage, root cause analysis, observability, incident ... Ensure outputs are traceable, testable, and auditable AI-Enabled DevSecOps, SDLC & AIOps * Embed AI ...

Senior AI Engineer

Reston, VA · On-site

$112K - $179K/yr

Integrate AI into IT operations (ticket triage, root cause analysis, observability, incident ... Ensure outputs are traceable, testable, and auditable AI-Enabled DevSecOps, SDLC & AIOps * Embed AI ...

Senior AI Engineer with Security Clearance

Reston, VA · On-site

$119K - $163K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Integrate AI into IT operations (ticket triage, root cause analysis, observability, incident ... Ensure outputs are traceable, testable, and auditableAI-Enabled DevSecOps, SDLC & AIOps * Embed AI ...

... observability. Leverage AI for log analysis, anomaly detection, root cause analysis, automation ... Experience or strong interest in applying AI in DevOps , including AIOps, intelligent monitoring ...

... observability. Leverage AI for log analysis, anomaly detection, root cause analysis, automation ... Experience or strong interest in applying AI in DevOps , including AIOps, intelligent monitoring ...

Director of DevOps

Herndon, VA · On-site

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

Stand up and continuously evolve an AIOps practice: AI-driven anomaly detection, log summarization ... Own vendor relationships and contracts for infrastructure, observability, and managed services ...

Experience with AIOps, observability platforms, and enterprise-scale monitoring strategies * Familiarity with DevSecOps pipelines, infrastructure-as-code, and platform engineering as enterprise ...

Showing results 41-60

Observability Aiops Engineer information

What are some common challenges faced by Observability AIOps engineers in integrating monitoring solutions across diverse technology stacks?

Observability AIOps Engineers often encounter challenges when integrating monitoring and analytics tools across a mix of legacy systems, cloud-native applications, and various third-party platforms. Ensuring consistent data collection, normalization, and visualization can be complex due to differing protocols, data formats, and tool compatibility. Collaboration with development, operations, and security teams is crucial to address these challenges, streamline workflows, and maintain a unified observability platform. Staying current with evolving AIOps technologies and best practices is also vital for continued success in this dynamic role.

What is an Observability AIOps engineer?

An Observability Aiops Engineer is a technology professional who focuses on implementing and managing observability tools and practices, often leveraging artificial intelligence for IT operations (AIOps). Their role is to ensure system reliability, performance, and uptime by monitoring, analyzing, and automating responses to IT incidents. They integrate data from logs, metrics, and traces to gain real-time insights, helping organizations quickly detect and resolve issues. This role combines expertise in software engineering, monitoring solutions, automation, and machine learning to improve the overall health and efficiency of IT environments.

What are the key skills and qualifications needed to thrive as an Observability AIOps engineer?

To thrive as an Observability AIOps Engineer, you need expertise in systems monitoring, data analytics, automation, and a strong understanding of IT infrastructure, often supported by a degree in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, Splunk, and AIOps platforms, as well as certifications in cloud solutions (AWS, Azure, or GCP), are typically required. Strong problem-solving skills, collaboration, and a proactive mindset help you stand out in identifying and addressing system anomalies. These skills and qualities are crucial for maintaining high system reliability, reducing downtime, and enabling data-driven decision-making in complex IT environments.

What is the difference between Observability Aiops Engineer vs Site Reliability Engineer?

AspectObservability Aiops EngineerSite Reliability Engineer
Primary FocusMonitoring, analyzing, and improving system observability using AI and automationEnsuring system reliability, scalability, and performance of services
Skills & CertificationsKnowledge of AI/ML, monitoring tools, scripting, cloud platformsSystems engineering, scripting, cloud infrastructure, incident management
Work EnvironmentDevOps teams, monitoring platforms, AI toolsOperations, development teams, cloud environments
Industry UsageTech companies, cloud providers, organizations focusing on AI-driven monitoringLarge-scale tech firms, SaaS providers, internet services

While both roles focus on system performance and reliability, the Observability Aiops Engineer specializes in leveraging AI and automation to enhance system observability, whereas the Site Reliability Engineer concentrates on maintaining overall system stability and scalability. Both roles often collaborate but have distinct core responsibilities.

What are popular job titles related to Observability Aiops Engineer jobs in Virginia?

For Observability Aiops Engineer jobs in Virginia, the most frequently searched job titles are:

What job categories do people searching Observability Aiops Engineer jobs in Virginia look for?

The top searched job categories for Observability Aiops Engineer jobs in Virginia are:

What cities in Virginia are hiring for Observability Aiops Engineer jobs?

Cities in Virginia with the most Observability Aiops Engineer job openings:

Full-time

Re-posted 10 days ago


Job description

Overview/ Job Responsibilities

Job Summary

We are seeking a skilled MLOps Engineer to join our team and ensure the seamless deployment, monitoring, and optimization of AI models in production.

The MLOps Engineer will design, implement, and maintain end-to-end machine learning pipelines, focusing on automating model deployment, monitoring model health, detecting data drift, and managing AI-related logging. This role will involve building scalable infrastructure and dashboards for real-time and historical insights, ensuring models are secure, performant, and aligned with business needs.

Key Responsibilities

  • Model Deployment: Deploy and manage machine learning models in production using tools like MLflow, Kubeflow, or AWS SageMaker, ensuring scalability and low latency.
  • Monitoring and Observability: Build and maintain dashboards using Grafana, Prometheus, or Kibana to track real-time model health (e.g., accuracy, latency) and historical trends.
  • Data Drift Detection: Implement drift detection pipelines using tools like Evidently AI or Alibi Detect to identify shifts in data distributions and trigger alerts or retraining.
  • Logging and Tracing: Set up centralized logging with ELK Stack or OpenTelemetry to capture AI inference events, errors, and audit trails for debugging and compliance.
  • Pipeline Automation: Develop CI/CD pipelines with GitHub Actions or Jenkins to automate model updates, testing, and deployment.
  • Security and Compliance: Apply secure-by-design principles to protect data pipelines and models, using encryption, access controls, and compliance with regulations like GDPR or NIST AI RMF.
  • Collaboration: Work with data scientists, AI Integration Engineers, and DevOps teams to align model performance with business requirements and infrastructure capabilities.
  • Optimization: Optimize models for production (e.g., via quantization or pruning) and ensure efficient resource usage on cloud platforms like AWS, Azure, or Google Cloud.
  • Documentation: Maintain clear documentation of pipelines, dashboards, and monitoring processes for cross-team transparency. 
Minimum Qualifications

Qualifications

  • Education: Bachelor's or Master's degree in Computer Science, Data Science, Engineering, or a related field.
  • Experience:
    • 5+ years in MLOps, DevOps, or software engineering with a focus on AI/ML systems.
    • Proven experience deploying models in production using MLflow, Kubeflow, or cloud platforms (AWS SageMaker, Azure ML).
    • Hands-on experience with observability tools like Prometheus, Grafana, or Datadog for real-time monitoring.
  • Technical Skills:
    • Proficiency in Python and SQL; familiarity with JavaScript or Go is a plus.
    • Expertise in containerization (Docker, Kubernetes) and CI/CD tools (GitHub Actions, Jenkins).
    • Knowledge of time-series databases (e.g., InfluxDB, TimescaleDB) and logging frameworks (e.g., ELK Stack, OpenTelemetry).
    • Experience with drift detection tools (e.g., Evidently AI, Alibi Detect) and visualization libraries (e.g., Plotly, Seaborn).
  • AI-Specific Skills:
    • Understanding of model performance metrics (e.g., precision, recall, AUC) and drift detection methods (e.g., KS test, PSI).
    • Familiarity with AI vulnerabilities (e.g., data poisoning, adversarial attacks) and mitigation tools like Adversarial Robustness Toolbox (ART).
  • Soft Skills:
    • Strong problem-solving and debugging skills for resolving pipeline and monitoring issues.
    • Excellent collaboration and communication skills to work with cross-functional teams.
    • Attention to detail for ensuring accurate and secure dashboard reporting.
  • Must be eligible to obtain a Department of Homeland Security EOD clearance ( Requirements 1. US Citizenship, 2. Favorable Background Investigation) 
Desired Qualifications

Preferred Qualifications

  • Experience with LLM monitoring tools like LangSmith or Helicone for generative AI applications.
  • Knowledge of compliance frameworks (e.g., GDPR, HIPAA) for secure data handling.
  • Contributions to open-source MLOps projects or familiarity with X platform discussions on #MLOps or #AIOps.
About Us

Formed through the strategic union of Sev1Tech and ERT, Entarian is a premier provider of mission-critical engineering and technology solutions. Founded on a legacy of excellence dating back to 1993, Entarian is a product of an evolved and fully diversified engineering and federal technology leader. From deep space to defense and civilian missions, Entarian delivers secure, mission-aligned digital solutions that drive national resilience and operational effectiveness. We don't just support modernization; we define it.

Join the Mission and Start your Career Journey: Apply Directly via our Careers Portal  Connect, Referrals & Inquiries? Email the team: careers@entarian.com

Entarian is an Equal Opportunity and Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, pregnancy, sexual orientation, gender identity, national origin, age, protected veteran status, or disability status.

Employment Type: OTHER