1

Observability Aiops Engineer Jobs in Texas (NOW HIRING)

AI with SRE

Irving, TX · On-site

$54.75 - $72.75/hr

Observability data collection and automation to help lead transformational initiatives within ... AIOPS Tools. * Practical experience implementing Golden Signals (latency, traffic, errors ...

New

DevOps Engineer

Dallas, TX · On-site

$52.25 - $71.50/hr

AI / MLOps / AIOps * Deploy, manage, and optimize AI/ML workloads in production environments ... Collaborate with engineering teams to improve observability, alerting, and operational readiness.

Implement observability, predictive monitoring, and AIOps capabilities to proactively prevent outages and improve reliability. Partner with Engineering, Infrastructure, Security, Architecture, and ...

Observability, AIOps, APM; Industry leading discovery technologies (SCCM, Tanium, Armis, Intune) and how they integrate with ServiceNow; Developing and re-engineering IT processes, capabilities, and ...

Observability, AIOps, APM; Industry leading discovery technologies (SCCM, Tanium, Armis, Intune) and how they integrate with ServiceNow; Developing and re-engineering IT processes, capabilities, and ...

Observability, AIOps, APM; Industry leading discovery technologies (SCCM, Tanium, Armis, Intune) and how they integrate with ServiceNow; Developing and re-engineering IT processes, capabilities, and ...

Observability, AIOps, APM; Industry leading discovery technologies (SCCM, Tanium, Armis, Intune) and how they integrate with ServiceNow; Developing and re-engineering IT processes, capabilities, and ...

Observability, AIOps, APM; Industry leading discovery technologies (SCCM, Tanium, Armis, Intune) and how they integrate with ServiceNow; Developing and re-engineering IT processes, capabilities, and ...

Senior DevOps Engineer

Dallas, TX · On-site

$128K - $165K/yr

AI / MLOps / AIOps * Deploy, manage, and optimize AI/ML workloads in production environments ... Collaborate with engineering teams to improve observability, alerting, and operational readiness.

Senior AI Quality & Reliability Engineer

Irving, TX · On-site

$84K - $114K/yr

... observability, reliability engineering, and modern AI Quality Engineering practices ... Partner with AI Engineering, AIOps, LLMOps, Security, Governance, Clinical, Data, and Product teams ...

... observability, predictive monitoring, and AIOps capabilities to proactively prevent outages and improve reliability. • Partner with Engineering, Infrastructure, Security, Architecture, and ...

... observability, predictive monitoring, and AIOps capabilities to proactively prevent outages and improve reliability. • Partner with Engineering, Infrastructure, Security, Architecture, and ...

Showing results 41-60

Observability Aiops Engineer information

What is an Observability AIOps engineer?

An Observability Aiops Engineer is a technology professional who focuses on implementing and managing observability tools and practices, often leveraging artificial intelligence for IT operations (AIOps). Their role is to ensure system reliability, performance, and uptime by monitoring, analyzing, and automating responses to IT incidents. They integrate data from logs, metrics, and traces to gain real-time insights, helping organizations quickly detect and resolve issues. This role combines expertise in software engineering, monitoring solutions, automation, and machine learning to improve the overall health and efficiency of IT environments.

What are the key skills and qualifications needed to thrive as an Observability AIOps engineer?

To thrive as an Observability AIOps Engineer, you need expertise in systems monitoring, data analytics, automation, and a strong understanding of IT infrastructure, often supported by a degree in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, Splunk, and AIOps platforms, as well as certifications in cloud solutions (AWS, Azure, or GCP), are typically required. Strong problem-solving skills, collaboration, and a proactive mindset help you stand out in identifying and addressing system anomalies. These skills and qualities are crucial for maintaining high system reliability, reducing downtime, and enabling data-driven decision-making in complex IT environments.

What are some common challenges faced by Observability AIOps engineers in integrating monitoring solutions across diverse technology stacks?

Observability AIOps Engineers often encounter challenges when integrating monitoring and analytics tools across a mix of legacy systems, cloud-native applications, and various third-party platforms. Ensuring consistent data collection, normalization, and visualization can be complex due to differing protocols, data formats, and tool compatibility. Collaboration with development, operations, and security teams is crucial to address these challenges, streamline workflows, and maintain a unified observability platform. Staying current with evolving AIOps technologies and best practices is also vital for continued success in this dynamic role.

What is the difference between Observability Aiops Engineer vs Site Reliability Engineer?

AspectObservability Aiops EngineerSite Reliability Engineer
Primary FocusMonitoring, analyzing, and improving system observability using AI and automationEnsuring system reliability, scalability, and performance of services
Skills & CertificationsKnowledge of AI/ML, monitoring tools, scripting, cloud platformsSystems engineering, scripting, cloud infrastructure, incident management
Work EnvironmentDevOps teams, monitoring platforms, AI toolsOperations, development teams, cloud environments
Industry UsageTech companies, cloud providers, organizations focusing on AI-driven monitoringLarge-scale tech firms, SaaS providers, internet services

While both roles focus on system performance and reliability, the Observability Aiops Engineer specializes in leveraging AI and automation to enhance system observability, whereas the Site Reliability Engineer concentrates on maintaining overall system stability and scalability. Both roles often collaborate but have distinct core responsibilities.

What job categories do people searching Observability Aiops Engineer jobs in Texas look for?

The top searched job categories for Observability Aiops Engineer jobs in Texas are:

What cities in Texas are hiring for Observability Aiops Engineer jobs?

Cities in Texas with the most Observability Aiops Engineer job openings:

Infographic showing various Observability Aiops Engineer job openings in Texas as of August 2026, with employment types broken down into 92% Full Time, 3% Part Time, and 5% Contract. Highlights an 86% Physical, 5% Hybrid, and 9% Remote job distribution.

$54.75 - $72.75/hr

Other

Posted 3 days ago

New


Job description

Hi 
 
Role  : SRE AI Engineer
Location : Austin TX
Job Description:
We are currently seeking a highly skilled SRE hands-on AI Engineer with solid experience in AI Observability and instrumentation approaches for AI systems, AI Agents development to perform detection, diagnosis and autonomous self-healing operates on AI Control Plane. 
Observability data collection and automation to help lead transformational initiatives within IT operations, encompassing development as well.  As a crucial figure in this role, you will participate/help with various technology domain groups and cross functional teams on unified observability gap analysis and solutioning (automation and manual fixes) 
Responsibilities:
  • Incorporate GenAI tooling and agentic capabilities to strengthen reliability outcomes across monitoring/alerting, rapid incident response, change management/testing, and DevOps/deployment processes. 
  • Experience building agentic workflows using LLMs, tool-calling, function-calling, multi-agent orchestration, and event-driven automation. 
  • Experience with Agent-to-Agent communication, AI agent federation, and enterprise AI control-plane concepts. 
  • Experience implementing AI control-plane governance, including policy-based execution, approval workflows, audit trails, guardrails, and risk-based remediation controls.
  • Expertise in Observability as a service, Dashboard as a services, monitoring as a services and alert as a service in all technology domains (application, infrastructure, database, security, middleware, network etc.,) Telemetry data collection using Dynatrace APM, SolarWinds, CISCO Switches, F5, Databases, Open-Source tools (Prometheus and Grafana), Log Aggregations (Kibana or Splunk) and AIOPS Tools.
  • Practical experience implementing Golden Signals (latency, traffic, errors, saturation) using related telemetry sources.
  • Configure application performance monitoring (APM), infrastructure monitoring, synthetic monitoring, RUM, and log monitoring.
  • Integrate Dynatrace with CI/CD pipelines, alerting tools, ITSM systems, and incident automation frameworks.
  • Tune alert thresholds, baselines, and AI-driven anomaly detection to reduce noise and improve actionable insights.
  • Deeper understanding of Login authentication mechanisms using Ping, ForgeRock and SiteMinder technologies (session management and cookie management) 
  • Define best practices and principles for SRE, including monitoring, alerting, and automation.
  • Collaborate with development teams on resiliency to ensure that services and applications are designed with operational reliability in mind.
  • Implement monitoring systems to assess the performance of applications and infrastructure and proactively identifying areas for optimization.
  • Ability to develop close relationship with other operational teams to integrate SRE practices and drive overall operational improvements across enterprise.
  • Stay up to date on industry trends, new technologies, and best practices in SRE and applying relevant advancements to the organization.
  • Ability to build strong working relationships across different levels, client focus mindset.
Qualifications:
  • Around 7-10 years of SRE hands on experience with AI OPS, cloud technologies, development, SRE toolsets and automation
  • Hands-on experience implementing Retrieval-Augmented Generation using approved enterprise knowledge sources such as runbooks, SOPs, RCA documents, incident history, architecture documents, and knowledge articles.
  • Expertise SPEC driven and Prompting using Ai IDE tools – Cursor, Kiro and Good Antigravity
  • Hands-on experience AI LLM’s – GPT - Open AI, Claude, Gemini etc.,
  • Experience with LangChain, LangGraph, Bedrock Agents, Azure AI Foundry, or Vertex AI Agent Builder 
  • Experience with vector databases such as OpenSearch, Pinecone, Chroma, Redis Vector, or pgvector.
  • Experience performing Observability current-state assessments, gap analysis and solutioning (automation and manual fixes) in all technology domains (application, infrastructure, database, security, middleware, network etc.,),
  • Strong hands-on automation experience in Observability as a code, dashboard as a code, monitoring as a code, alert as a code (Instrumentation, templates, automatic deployment, visualization and alerting)
  • Strong hands-on experience with any Cloud Technology (AWS): Control Tower, Project Setup, Creating Accounts, RDS, SSO
  • Monitoring & alerting setup experience with Splunk, Prometheus, Grafana, Kibana, ELK, with pref. for APM (Dynatrace).
  • Strong skills in APM, distributed tracing, synthetic & real user monitoring, log monitoring, and Davis AI configuration
  • Own the design, configuration, CICD deployment, and optimization for enterprise-wide observability tools.
  • Experience integrating, automation, and cloud platforms (AWS, Azure, Google Cloud Platform).
  • Extended experience instrumenting OTEL Framework.
  • Hands on experience with Dynatrace Plug-and-play observability modules (OKit) development for Observability Developers Java and .Net applications.
  • Define monitoring standards, best practices, and governance to ensure consistency and scalability.
  • Experience to deploy and tune OneAgent, build end-to-end PurePath tracing, and leverage Smartscape topology for proactive performance monitoring and root-cause analysis.
  • Collaborate with application and infrastructure teams to troubleshoot performance issues and implement permanent fixes.
Good to have:
  • Any of the relevant professional certifications – AIOPS related certifications, Certified Site Reliability Engineer (CSRE), Certified Kubernetes Administrator (CKA), AWS Certified DevOps Engineer Professional, , Google Cloud Professional; DevOps Engineer