You will define and drive the AIOps strategy for the organization, combining cloud operations, observability, automation, SRE practices, and AI/agentic solutions to improve reliability, incident ...
You will define and drive the AIOps strategy for the organization, combining cloud operations, observability, automation, SRE practices, and AI/agentic solutions to improve reliability, incident ...
AVP, Enterprise Database Platforms & Operations
Atlanta, GA · Hybrid
$47.75 - $65.75/hr
Advance AIOps capabilities for predictive insights and proactive incident detection * Drive measurable reduction in incident volume and MTTR through automation and observability integration
AVP, Enterprise Database Platforms & Operations
Atlanta, GA · Hybrid
$47.75 - $65.75/hr
Advance AIOps capabilities for predictive insights and proactive incident detection * Drive measurable reduction in incident volume and MTTR through automation and observability integration
AWS DevOps Datadog
Sandy Springs, GA · On-site
$52.25 - $71.75/hr
NTT DATA''s Client is currently seeking an AWS DevOps / SRE Engineer - Datadog & AIOps Location: Atlanta, Georgia -- Preferred Onsite/Hybrid Employment Type: Full-Time Experience Level: 5+ years in ...
AWS DevOps Datadog
Sandy Springs, GA · On-site
$52.25 - $71.75/hr
NTT DATA''s Client is currently seeking an AWS DevOps / SRE Engineer - Datadog & AIOps Location: Atlanta, Georgia -- Preferred Onsite/Hybrid Employment Type: Full-Time Experience Level: 5+ years in ...
This role will serve as the technical authority for log management, observability engineering, telemetry pipelines, AIOps, security analytics, and data lakehouse architectures leveraging Splunk ...
This role will serve as the technical authority for log management, observability engineering, telemetry pipelines, AIOps, security analytics, and data lakehouse architectures leveraging Splunk ...
This role will serve as the technical authority for log management, observability engineering, telemetry pipelines, AIOps, security analytics, and data lakehouse architectures leveraging Splunk ...
This role will serve as the technical authority for log management, observability engineering, telemetry pipelines, AIOps, security analytics, and data lakehouse architectures leveraging Splunk ...
Deploy and manage AIOps monitoring for predictive detection, automated anomaly identification, and AI-driven incident response. * Architect technical strategy for critical applications - from roadmap ...
Deploy and manage AIOps monitoring for predictive detection, automated anomaly identification, and AI-driven incident response. * Architect technical strategy for critical applications - from roadmap ...
Senior Incident Management Practice Engineer
Atlanta, GA · On-site +1
$110K - $151K/yr
This role drives operational resilience by leveraging AIOps, advanced analytics, and standardized processes across ITSM functions to proactively detect, respond to, and prevent service disruptions ...
Senior Incident Management Practice Engineer
Atlanta, GA · On-site +1
$110K - $151K/yr
This role drives operational resilience by leveraging AIOps, advanced analytics, and standardized processes across ITSM functions to proactively detect, respond to, and prevent service disruptions ...
Senior Incident Management Practice Engineer
Atlanta, GA · On-site +1
$110K - $151K/yr
This role drives operational resilience by leveraging AIOps, advanced analytics, and standardized processes across ITSM functions to proactively detect, respond to, and prevent service disruptions ...
Senior Incident Management Practice Engineer
Atlanta, GA · On-site +1
$110K - $151K/yr
This role drives operational resilience by leveraging AIOps, advanced analytics, and standardized processes across ITSM functions to proactively detect, respond to, and prevent service disruptions ...
Senior Incident Management Practice Engineer
Atlanta, GA · On-site +1
$110K - $151K/yr
This role drives operational resilience by leveraging AIOps, advanced analytics, and standardized processes across ITSM functions to proactively detect, respond to, and prevent service disruptions ...
Senior Incident Management Practice Engineer
Atlanta, GA · On-site +1
$110K - $151K/yr
This role drives operational resilience by leveraging AIOps, advanced analytics, and standardized processes across ITSM functions to proactively detect, respond to, and prevent service disruptions ...
AI & Automation Enablement Champion adoption of AI-powered capabilities within ServiceNow (e.g., Virtual Agent, Predictive Intelligence, AIOps). Identify opportunities to integrate AI/ML solutions to ...
Quick apply
AI & Automation Enablement Champion adoption of AI-powered capabilities within ServiceNow (e.g., Virtual Agent, Predictive Intelligence, AIOps). Identify opportunities to integrate AI/ML solutions to ...
Dynatrace Observability consultant CL LP
Atlanta, GA · On-site
$54.75 - $72.75/hr
Collaborate with teams to optimize Dynatrace usage for AIOps-driven insights and automated anomaly detection. Provide oversight for production operations to maximize reliability and automation.
Dynatrace Observability consultant CL LP
Atlanta, GA · On-site
$54.75 - $72.75/hr
Collaborate with teams to optimize Dynatrace usage for AIOps-driven insights and automated anomaly detection. Provide oversight for production operations to maximize reliability and automation.
Senior Systems Engineer
Atlanta, GA · On-site
$100K - $137K/yr
Experience with observability or AIOps platforms. * Operations or administration experience with specialty business systems (e.g., IBM iSeries / AS400). * Experience operating within regulated ...
Senior Systems Engineer
Atlanta, GA · On-site
$100K - $137K/yr
Experience with observability or AIOps platforms. * Operations or administration experience with specialty business systems (e.g., IBM iSeries / AS400). * Experience operating within regulated ...
AIOps & agentic operations (Preferred): demonstratedexperience applying AI/ML and GenAI to operations, including anomaly detection, event correlation, automated remediation, and building or operating ...
New
AIOps & agentic operations (Preferred): demonstratedexperience applying AI/ML and GenAI to operations, including anomaly detection, event correlation, automated remediation, and building or operating ...
New
Principal Infrastructue Architect
Atlanta, GA · On-site
$85 - $90/hr
Identify and implement opportunities to leverage emerging technologies, AI/AIOps, and intelligent automation to improve efficiency and service delivery * Provide architectural oversight for major ...
Principal Infrastructue Architect
Atlanta, GA · On-site
$85 - $90/hr
Identify and implement opportunities to leverage emerging technologies, AI/AIOps, and intelligent automation to improve efficiency and service delivery * Provide architectural oversight for major ...
AIOps & agentic operations (Preferred): demonstrated experience applying AI/ML and GenAI to operations, including anomaly detection, event correlation, automated remediation, and building or ...
AIOps & agentic operations (Preferred): demonstrated experience applying AI/ML and GenAI to operations, including anomaly detection, event correlation, automated remediation, and building or ...
Jr. - Mid Level DevOps Engineer
$50.75 - $69.50/hr
E xperience with AIOps or predictive automation. * L eadership or mentoring experience. What You'll Gain: * C ompetitive salary and benefits * F lexible work environment * L earning and certification ...
Jr. - Mid Level DevOps Engineer
$50.75 - $69.50/hr
E xperience with AIOps or predictive automation. * L eadership or mentoring experience. What You'll Gain: * C ompetitive salary and benefits * F lexible work environment * L earning and certification ...
VCF Engineer with Security Clearance
Warner Robins, GA · On-site
$100K - $140K/yr
Integrating Key Performance Indicators (KPIs) with ITSI features Enriching the environment with machine learning and AIOps Conducting adoption assessments Develop response architectures and provide ...
VCF Engineer with Security Clearance
Warner Robins, GA · On-site
$100K - $140K/yr
Integrating Key Performance Indicators (KPIs) with ITSI features Enriching the environment with machine learning and AIOps Conducting adoption assessments Develop response architectures and provide ...
Senior Product Marketing Manager - Observability
Alpharetta, GA · Hybrid
$118K - $154K/yr
Build thought leadership about observability trends, hybrid cloud operations, and AIOps through analyst engagement, speaking opportunities, and industry content. * Market Research and Competitive ...
Posted today
Senior Product Marketing Manager - Observability
Alpharetta, GA · Hybrid
$118K - $154K/yr
Build thought leadership about observability trends, hybrid cloud operations, and AIOps through analyst engagement, speaking opportunities, and industry content. * Market Research and Competitive ...
Posted today
Experience with AIOps, observability (Datadog, Dynatrace, Splunk, New Relic), and intelligent automation platforms * Working knowledge of cloud security, Zero Trust architecture, compliance ...
Experience with AIOps, observability (Datadog, Dynatrace, Splunk, New Relic), and intelligent automation platforms * Working knowledge of cloud security, Zero Trust architecture, compliance ...
Experience with AIOps, observability (Datadog, Dynatrace, Splunk, New Relic), and intelligent automation platforms * Working knowledge of cloud security, Zero Trust architecture, compliance ...
Experience with AIOps, observability (Datadog, Dynatrace, Splunk, New Relic), and intelligent automation platforms * Working knowledge of cloud security, Zero Trust architecture, compliance ...
Aiops information
What does an AIOps do?
In an AIOps role, your primary responsibilities include monitoring IT infrastructure, analyzing large volumes of system data, and proactively automating responses to incidents using machine learning models and advanced analytics. You will collaborate closely with IT, DevOps, and cybersecurity teams to detect anomalies, optimize system performance, and reduce downtime. Troubleshooting and refining automation scripts are part of the daily workflow, along with participating in incident response and post-mortem analysis. This role requires continual learning, as you will regularly implement new tools and processes to stay ahead of emerging technology trends and operational challenges.
What are the key skills and qualifications needed to thrive in an AIOps?
To thrive in an AIOps role, you need a strong background in IT operations, data analytics, and automation technology, often supported by a degree in computer science or a related field. Familiarity with monitoring tools like Splunk, ELK Stack, and AI/ML platforms, as well as certifications in cloud platforms or DevOps practices, is highly valuable. Excellent problem-solving skills, adaptability, and effective communication are essential for collaborating with cross-functional teams. These competencies enable AIOps professionals to effectively predict, identify, and resolve IT issues rapidly, ensuring seamless and efficient system operations.
What is an AIOps?
An AIOps job involves using artificial intelligence and machine learning to enhance IT operations by automating processes, analyzing vast amounts of data, and identifying patterns to prevent issues. AIOps professionals work with tools that help in real-time monitoring, anomaly detection, and incident response to improve system reliability and efficiency. They collaborate with IT teams to reduce downtime, improve performance, and streamline operations.
How to become an Aiops engineer?

Full-time
Medical, Dental, Vision, Life, Retirement, PTO
Re-posted 15 days ago
Zelis rating
7.6
Based on 11 frontline employees who took The Breakroom Quiz
153rd of 242 rated software companies
Job description
At Zelis, we Get Stuff Done. So, let's get to it!
A Little About Us
Zelis is modernizing the healthcare financial experience across payers, providers, and healthcare consumers. We serve more than 750 payers, including the top five national health plans, regional health plans, TPAs and millions of healthcare providers and consumers across our platform of solutions. Zelis sees across the system to identify, optimize, and solve problems holistically with technology built by healthcare experts - driving real, measurable results for clients.
At Zelis, AI is woven into the fabric of how we work. Every associate is expected - and empowered - to partner with AI to challenge the status quo, accelerate innovation, and amplify their impact. This is a place for builders with a growth mindset who act with agility, embrace change, and use modern technology to shape smarter solutions, exceptional experiences, and the future of our industry for our clients, customers, and our culture.
A Little About You
You bring a unique blend of personality and professional expertise to your work, inspiring others with your passion and dedication. Your career is a testament to your diverse experiences, community involvement, and the valuable lessons you've learned along the way. You are more than just your resume; you are a reflection of your achievements, the knowledge you've gained, and the personal interests that shape who you are.
Position Overview
This role will lead the next phase of Zelis' operational transformation as we accelerate AWS migration and expand AI-native capabilities into the operations space. You will define and drive the AIOps strategy for the organization, combining cloud operations, observability, automation, SRE practices, and AI/agentic solutions to improve reliability, incident response, operational efficiency, and platform resilience.As a senior technical leader, you will work across Engineering, Cloud, Infrastructure, Security, and Operations teams to design intelligent operational capabilities that move beyond traditional monitoring into proactive, automated, and agent-assisted operations. This role requires strong AWS depth, hands-on experience with observability platforms such as New Relic and OpenSearch, strong AI and agent experience, a strong SRE mindset, and proven ability to apply ChatOps and agentic or multi-agent systems to real-world operational workflows.What You'll Do
Lead the AIOps strategy and architecture for Zelis Price Business Unit as we modernize operations alongside AWS migration and AI-native acceleration.
Define and implement intelligent operational patterns that improve incident detection, triage, remediation, root cause analysis, and operational decision-making.
Architect and scale observability capabilities across cloud and application environments, including metrics, logs, traces, dashboards, alerting, and service health visibility.
Drive operational excellence across AWS environments by establishing scalable patterns for monitoring, resilience, reliability, automation, and governance.
Design and implement AIOps capabilities using AI, agents, and agentic workflows to support incident response, anomaly detection, alert correlation, noise reduction, troubleshooting, and operational automation.
Lead development of agentic and multi-agent operational solutions that coordinate across monitoring, diagnostics, knowledge retrieval, remediation workflows, and operator assistance.
Build and mature ChatOps capabilities to improve collaboration, visibility, response speed, and workflow automation across engineering and operations teams.
Partner with SRE, Cloud Engineering, Infrastructure, Security, and Application teams to embed reliability engineering practices into operational processes and platform design.
Establish standards for operational telemetry, service-level objectives, alert quality, escalation workflows, incident readiness, and post-incident learning.
Drive adoption and optimization of observability tools such as New Relic, OpenSearch, and related monitoring, logging, and analytics platforms.
Identify opportunities to apply AI to reduce manual operational effort, improve mean time to detect and resolve issues, and increase platform stability and operator productivity.
Ensure AIOps solutions are implemented with strong governance, security, auditability, and operational trustworthiness.
Create playbooks, standards, reusable patterns, and operating models that scale AIOps adoption across teams.
Mentor engineers and operators in modern operations practices spanning observability, automation, SRE, ChatOps, and AI-assisted operations.
What You'll Bring to ZelisExperience & leadership: Typically BS + 12 years or MS + 10 years (or equivalent), with a strong track record leading cloud operations, platform operations, SRE, observability, or AIOps initiatives across complex enterprise environments.
AWS depth: Strong hands-on experience designing and operating workloads on AWS, with expertise across compute, networking, storage, security, automation, and cloud operations patterns.
Observability expertise: Deep experience with modern observability and monitoring platforms such as New Relic, OpenSearch, and related tools for metrics, logs, traces, dashboards, alerting, and operational analytics.
AIOps experience: Proven experience applying AI to operations use cases such as event correlation, anomaly detection, alert reduction, root cause analysis, remediation support, and operational workflow automation.
Agentic AI fluency: Strong experience designing or implementing AI agents, agentic workflows, or multi-agent systems that improve operational processes and operator effectiveness.
SRE mindset: Strong grounding in site reliability engineering principles, including service reliability, SLOs/SLIs, error budgets, automation, incident management, resilience, and continuous improvement.
ChatOps experience: Demonstrated success building or scaling ChatOps practices that improve collaboration, incident response, and operational execution through integrated messaging and workflow automation.
Automation & engineering strength: Strong knowledge of scripting, infrastructure automation, operational tooling, APIs, event-driven systems, and platform integration patterns.
Problem-solving & execution: Ability to translate operational pain points into scalable technical solutions that improve reliability, speed, and operational maturity.
Communication & influence: Able to influence technical teams and senior leaders, build alignment across functions, and communicate complex operational strategies clearly.
Governance & trust: Experience implementing operational AI responsibly with appropriate controls for accuracy, security, compliance, explainability, and human oversight.
Preferred Qualifications
Experience leading AIOps or intelligent operations initiatives in a cloud-first or large-scale enterprise environment.
Experience supporting AWS migration programs and modern cloud operating models.
Familiarity with incident management tooling, runbook automation, knowledge systems, and operational workflow orchestration platforms.
Experience integrating AI agents with observability, ticketing, collaboration, or operational systems.
Experience in healthcare, regulated environments, or other domains requiring strong reliability and compliance practices.
Exposure to platform engineering, DevOps, and developer experience practices that intersect with operational excellence.
Please note at this time we are unable to proceed with candidates who require visa sponsorship now or in the future.
Location and Workplace Flexibility
Zelis is headquartered in the U.S., with multiple locations across the country and in Hyderabad, India. Check out our locations to learn more about our offices. All employee work locations are based on the needs of the position and are determined by the Leadership team. In-office work and activities vary based on work and team objectives in accordance with Company policies.
While location expectations vary by role, candidates within approximately 50 miles of a U.S. office are generally preferred to support collaboration when needed. Our hybrid approach is flexible, and in-office presence is guided by team and business needs rather than a fixed weekly schedule.
Base Salary Range
$169,000.00 - $213,750.00At Zelis we are committed to providing fair and equitable compensation packages. The base salary range allows us to make an offer that considers multiple individualized factors, including experience, education, qualifications, as well as job-related and industry-related knowledge and skills, etc. Base pay is just one part of our Total Rewards package, which may also include discretionary bonus plans, commissions, or other incentives depending on the role.
Zelis' full-time associates are eligible for a highly competitive benefits package as well, which demonstrates our commitment to our employees' health, well-being, and financial protection. The US-based benefits include a 401k plan with employer match, flexible paid time off, holidays, parental leaves, life and disability insurance, and health benefits including medical, dental, vision, and prescription drug coverage.
Equal Employment Opportunity
Zelis is proud to be an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.
We welcome applicants from all backgrounds and encourage you to apply even if you don't meet 100% of the qualifications for the role. We believe in the value of diverse perspectives and experiences and are committed to building an inclusive workplace for all.
Accessibility Support
We are dedicated to ensuring our application process is accessible to all candidates. If you are a qualified individual with a disability or a disabled veteran and require a reasonable accommodation with any part of the application and/or interview process, please email TalentAcquisition@zelis.com.
Disclaimer
The above statements are intended to describe the general nature and level of work being performed by people assigned to this classification. They are not to be construed as an exhaustive list of all responsibilities, duties, and skills required of personnel so classified. All personnel may be required to perform duties outside of their normal responsibilities, duties, and skills from time to time.
About Zelis
Sourced by ZipRecruiter
Industry
Software development
Company size
1,001 - 5,000 Employees
Headquarters location
Bedminster, NJ, US
Year founded
2016