1

Observability Aiops Engineer Jobs in Texas (NOW HIRING)

You will define and drive the AIOps strategy for the organization, combining cloud operations, observability, automation, SRE practices, and AI/agentic solutions to improve reliability, incident ...

Enterprise Observability Architect

Coppell, TX ยท On-site

$64.25 - $82.75/hr

... AIOps-enablement * Establish enterprise-wide architectural standards, patterns, and controls for ... Drive platform-engineering approaches that deliver observability as a scalable, self-service ...

New

DevOps Engineer

Austin, TX

$52.25 - $71.50/hr

I mplement monitoring, logging, and observability solutions (Prometheus, Grafana, ELK, Datadog ... E xperience with AIOps or predictive automation. * L eadership or mentoring experience. What You'll ...

DevOps Engineer

Addison, TX ยท On-site

$51 - $70/hr

... observability implementation. - Networking and security fundamentals. Preferred Qualifications: * - AWS/Azure/GCP DevOps certifications. - GitOps tools such as ArgoCD or Flux. - Experience with AIOps ...

next page

Showing results 1-20

Observability Aiops Engineer information

What are some common challenges faced by Observability AIOps Engineers in integrating monitoring solutions across diverse technology stacks?

Observability AIOps Engineers often encounter challenges when integrating monitoring and analytics tools across a mix of legacy systems, cloud-native applications, and various third-party platforms. Ensuring consistent data collection, normalization, and visualization can be complex due to differing protocols, data formats, and tool compatibility. Collaboration with development, operations, and security teams is crucial to address these challenges, streamline workflows, and maintain a unified observability platform. Staying current with evolving AIOps technologies and best practices is also vital for continued success in this dynamic role.

What is an Observability Aiops Engineer?

An Observability Aiops Engineer is a technology professional who focuses on implementing and managing observability tools and practices, often leveraging artificial intelligence for IT operations (AIOps). Their role is to ensure system reliability, performance, and uptime by monitoring, analyzing, and automating responses to IT incidents. They integrate data from logs, metrics, and traces to gain real-time insights, helping organizations quickly detect and resolve issues. This role combines expertise in software engineering, monitoring solutions, automation, and machine learning to improve the overall health and efficiency of IT environments.

What are the key skills and qualifications needed to thrive as an Observability AIOps Engineer, and why are they important?

To thrive as an Observability AIOps Engineer, you need expertise in systems monitoring, data analytics, automation, and a strong understanding of IT infrastructure, often supported by a degree in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, Splunk, and AIOps platforms, as well as certifications in cloud solutions (AWS, Azure, or GCP), are typically required. Strong problem-solving skills, collaboration, and a proactive mindset help you stand out in identifying and addressing system anomalies. These skills and qualities are crucial for maintaining high system reliability, reducing downtime, and enabling data-driven decision-making in complex IT environments.

What is the difference between Observability Aiops Engineer vs Site Reliability Engineer?

AspectObservability Aiops EngineerSite Reliability Engineer
Primary FocusMonitoring, analyzing, and improving system observability using AI and automationEnsuring system reliability, scalability, and performance of services
Skills & CertificationsKnowledge of AI/ML, monitoring tools, scripting, cloud platformsSystems engineering, scripting, cloud infrastructure, incident management
Work EnvironmentDevOps teams, monitoring platforms, AI toolsOperations, development teams, cloud environments
Industry UsageTech companies, cloud providers, organizations focusing on AI-driven monitoringLarge-scale tech firms, SaaS providers, internet services

While both roles focus on system performance and reliability, the Observability Aiops Engineer specializes in leveraging AI and automation to enhance system observability, whereas the Site Reliability Engineer concentrates on maintaining overall system stability and scalability. Both roles often collaborate but have distinct core responsibilities.

What are popular job titles related to Observability Aiops Engineer jobs in Texas? For Observability Aiops Engineer jobs in Texas, the most frequently searched job titles are:
What job categories do people searching Observability Aiops Engineer jobs in Texas look for? The top searched job categories for Observability Aiops Engineer jobs in Texas are:
What cities in Texas are hiring for Observability Aiops Engineer jobs? Cities in Texas with the most Observability Aiops Engineer job openings:
Infographic showing various Observability Aiops Engineer job openings in Texas as of July 2026, with employment types broken down into 7% As Needed, 61% Full Time, 3% Part Time, 3% Contract, and 26% Nights. Highlights an 77% Physical, 8% Hybrid, and 15% Remote job distribution.

$54.75 - $72.75/hr

Full-time

Posted 18 days ago


Job description

Observability Engineer | Irving, Texas, United States Job Summary: Senior Observability Engineer - Irving, TX (Onsite) About the Role Join a dynamic technology team as a Senior Observability Engineer based in Irving, TX. You will architect and lead end-to-end observability solutions across complex, hybrid environments, transforming platform telemetry into actionable insights that drive reliability and performance. This is an opportunity to shape observability strategy, work with cutting-edge tools, and collaborate with engineering, operations, and leadership. Advance your career by building and standardizing monitoring frameworks at scale. Responsibilities - Architect and implement comprehensive observability frameworks across cloud, on-premises, networking, databases, middleware, and applications - Evaluate, select, and integrate observability and monitoring tools, establishing reference architectures - Design and maintain standardized Grafana dashboards for platform and workload health (OCP, AKS, GKE) - Define golden signals and platform health KPIs tied to availability, performance, and reliability - Serve as an advanced Splunk user: develop complex SPL queries, dashboards, and root-cause investigations - Correlate logs, metrics, and events across Grafana and Splunk to drive rapid incident resolution (MTTR reduction) - Implement and tune platform-specific observability for Kubernetes platforms (OCP, AKS, GKE) - Configure and manage ThousandEyes for synthetic monitoring and network intelligence - Administer BigPanda for AIOps-driven event correlation and noise reduction - Integrate ServiceNow for automated incident creation and enriched alerting - Document observability standards, dashboards, and onboarding processes Required Skills and Experience - 7+ years in IT operations, SRE, systems engineering, or infrastructure - 5+ years designing and implementing observability/monitoring solutions across distributed systems - 5+ years production experience with Kubernetes platforms - Expertise in logging, tracing, and enhanced monitoring - Strong hands-on experience with Grafana (dashboards, alerts, data sources) - Advanced Splunk SPL, dashboarding, and investigation capabilities - Proficiency with querying languages (SQL, PromQL) - Deep understanding of Kubernetes internals, OpenShift (OCP), AKS, GKE - Experience with Prometheus, OpenTelemetry, Kubernetes exporters - OS-level monitoring (Linux/Windows) and network fundamentals - Experience with ThousandEyes, BigPanda, and ServiceNow ITSM workflows Preferred Skills - Experience designing SLOs/SLIs, reliability scorecards - Familiarity with Istio, service mesh metrics, mTLS - Capacity planning and trend analysis using observability data - Exposure to multi-cloud observability strategies - Monitoring for databases, message brokers, middleware - Familiarity with AIOps or ML-driven anomaly detection Benefits - Work in a collaborative, technology-driven environment - Direct impact on reliability, performance, and operational excellence - Exposure to the latest observability and cloud-native technologies - Career growth opportunities in a high-visibility engineering role How to Apply Ready to drive observability excellence and elevate platform reliability? Submit your resume today to join our Irving, TX team and advance your career as a Senior Observability Engineer.

Indotronix logo

About Indotronix

Sourced by ZipRecruiter

In 1986, Indotronix established itself in the staffing space. 22 years later, Avani entered the scene, offering consulting and technology development. Finally, in 2016, the two joined forces to begin delivering talent across all areas, from Staffing to Consulting to unique platform development.

Industry

Recruiting and staffing services

Company size

1,001 - 5,000 Employees

Headquarters location

Rochester, NY, US