1

Observability Aiops Engineer Jobs (NOW HIRING)

Senior AIOps ML Engineer

Los Angeles, CA · On-site

$112K - $154K/yr

The engineer will also focus on security and compliance observability, collaborating with security ... ML Model Development & AIOps: Design, train, and deploy machine learning models for streaming ...

AI OPS Engineer

Fort Belvoir, VA · On-site

$160 - $175/hr

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Engineer secure data ingestion pipelines for telemetry data from cross-domain solutions and ...

Senior Site Reliability Engineer, AIOPs

Santa Clara, CA · On-site

$67 - $89/hr

Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing ... Proven programming experience building automation tools or services - ideally in Python, or similar ...

Senior AI OPS Engineer

Fort Belvoir, VA · On-site

$150K - $174K/yr

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Engineer secure data ingestion pipelines for telemetry data from cross-domain solutions and ...

Senior AI OPS Engineer

Fort Belvoir, VA · On-site

$150K - $174K/yr

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Engineer secure data ingestion pipelines for telemetry data from cross-domain solutions and ...

$103K - $155K/yr

Mission-Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Engineer secure data ingestion pipelines for telemetry data from cross-domain solutions and ...

NY · On-site

$180 - $260/hr

Strong track record leading cloud operations, platform operations, SRE, observability, or AIOps initiatives across complex enterprise environments * Strong hands‑on experience designing and ...

Senior AI OPS Engineer

Albuquerque, NM · On-site

$103 - $155/hr

Mission‑Critical Observability: Architect and maintain Splunk AIOps solutions across unclassified ... Engineer secure data ingestion pipelines for telemetry data from cross-domain solutions and ...

Showing results 21-40

Observability Aiops Engineer information

What is an Observability AIOps engineer?

An Observability Aiops Engineer is a technology professional who focuses on implementing and managing observability tools and practices, often leveraging artificial intelligence for IT operations (AIOps). Their role is to ensure system reliability, performance, and uptime by monitoring, analyzing, and automating responses to IT incidents. They integrate data from logs, metrics, and traces to gain real-time insights, helping organizations quickly detect and resolve issues. This role combines expertise in software engineering, monitoring solutions, automation, and machine learning to improve the overall health and efficiency of IT environments.

What are the key skills and qualifications needed to thrive as an Observability AIOps engineer?

To thrive as an Observability AIOps Engineer, you need expertise in systems monitoring, data analytics, automation, and a strong understanding of IT infrastructure, often supported by a degree in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, Splunk, and AIOps platforms, as well as certifications in cloud solutions (AWS, Azure, or GCP), are typically required. Strong problem-solving skills, collaboration, and a proactive mindset help you stand out in identifying and addressing system anomalies. These skills and qualities are crucial for maintaining high system reliability, reducing downtime, and enabling data-driven decision-making in complex IT environments.

What are some common challenges faced by Observability AIOps engineers in integrating monitoring solutions across diverse technology stacks?

Observability AIOps Engineers often encounter challenges when integrating monitoring and analytics tools across a mix of legacy systems, cloud-native applications, and various third-party platforms. Ensuring consistent data collection, normalization, and visualization can be complex due to differing protocols, data formats, and tool compatibility. Collaboration with development, operations, and security teams is crucial to address these challenges, streamline workflows, and maintain a unified observability platform. Staying current with evolving AIOps technologies and best practices is also vital for continued success in this dynamic role.

What is the difference between Observability Aiops Engineer vs Site Reliability Engineer?

AspectObservability Aiops EngineerSite Reliability Engineer
Primary FocusMonitoring, analyzing, and improving system observability using AI and automationEnsuring system reliability, scalability, and performance of services
Skills & CertificationsKnowledge of AI/ML, monitoring tools, scripting, cloud platformsSystems engineering, scripting, cloud infrastructure, incident management
Work EnvironmentDevOps teams, monitoring platforms, AI toolsOperations, development teams, cloud environments
Industry UsageTech companies, cloud providers, organizations focusing on AI-driven monitoringLarge-scale tech firms, SaaS providers, internet services

While both roles focus on system performance and reliability, the Observability Aiops Engineer specializes in leveraging AI and automation to enhance system observability, whereas the Site Reliability Engineer concentrates on maintaining overall system stability and scalability. Both roles often collaborate but have distinct core responsibilities.

More about Observability Aiops Engineer jobs

What cities are hiring for Observability Aiops Engineer jobs?

Cities with the most Observability Aiops Engineer job openings:

What states have the most Observability Aiops Engineer jobs?

States with the most job openings for Observability Aiops Engineer jobs include:

What job categories do people searching Observability Aiops Engineer jobs look for?

The top searched job categories for Observability Aiops Engineer jobs are:

Infographic showing various Observability Aiops Engineer job openings in the United States as of August 2026, with employment types broken down into 94% Full Time, 2% Part Time, and 4% Contract. Highlights an 85% Physical, 6% Hybrid, and 9% Remote job distribution.

Senior IA Ops Engineer with Security Clearance

Tetrad Digital Integrity (TDI)

Fort Belvoir, VA • On-site

$129K - $177K/yr

Other

Re-posted 23 days ago


Job description

Tetrad Digital Integrity (TDI) is a leading-edge cybersecurity firm with a mission to safeguard and protect our customers from increasing threats and vulnerabilities in this digital age.  TDI is seeking a Senior AIOps Engineer to lead ITSM transformation efforts within a secure mission environment. As the technical lead for this initiative, you will orchestrate integrations across existing Network Engineering, ServiceNow, and SolarWinds teams. Utilize Splunk and Machine Learning Toolkit to provide descriptive and predictive analytics and establish closed-loop automated incident response, ensuring the high availability of mission-essential infrastructure. This position requires fully onsite support to the Fort Belvoir, VA area and active TS/SCI level clearance.  RESPONSIBILITIES:
Lead AIOps platform integration efforts across Network Engineering, ServiceNow, and SolarWinds teams to establish unified observability and telemetry capabilities.
Architect and maintain Splunk AIOps and ITSI solutions across classified and unclassified environments, delivering real-time situational awareness, event correlation, and automated incident remediation through ServiceNow integration.
Develop and deploy advanced analytics and machine learning models using Splunk MLTK to detect anomalies, identify cyber threats, predict infrastructure issues, and reduce alert fatigue.
Engineer secure telemetry ingestion and correlation pipelines from enterprise infrastructure, cross-domain solutions, and tactical edge systems to provide a comprehensive view of operational health.
Support defensive cyber operations by integrating AIOps insights into security workflows, while ensuring compliance with DoD STIGs, IL5/IL6 requirements, and maintaining technical and architectural documentation.
QUALIFICATIONS:
Active TS/SCI security clearance
Candidates must possess DoD IAT Level II certification (e.g., Security+ CE, CySA+, GSEC, or SSCP)
Bachelor's degree and 7+ years of Splunk Enterprise experience, including architecture, cluster administration, and advanced SPL development.
3+ years of experience implementing AIOps workflows and integrating Splunk with ServiceNow or other enterprise ITSM platforms.
Experience building, tuning, and deploying machine learning models using Splunk MLTK.
Strong scripting and automation skills, including Python, API integrations, custom search commands, and automated remediation solutions.
Must be able to present designs, plans, and analyses of alternatives to technical leadership boards for approvals. PREFERRED QUALIFICATIONS:
Splunk Enterprise Certified Architect or Splunk ITSI Certified Admin.
 Experience with Cloud Native Computing Foundation (CNCF) observability tools in secure hybrid multi-cloud environments (Azure/AWS).