1

Observability Aiops Engineer Jobs in Ohio (NOW HIRING)

Observability, AIOps, APM; Industry leading discovery technologies (SCCM, Tanium, Armis, Intune) and how they integrate with ServiceNow; Developing and re-engineering IT processes, capabilities, and ...

Observability, AIOps, APM; Industry leading discovery technologies (SCCM, Tanium, Armis, Intune) and how they integrate with ServiceNow; Developing and re-engineering IT processes, capabilities, and ...

next page

Showing results 1-20

Observability Aiops Engineer information

What is an Observability AIOps engineer?

An Observability Aiops Engineer is a technology professional who focuses on implementing and managing observability tools and practices, often leveraging artificial intelligence for IT operations (AIOps). Their role is to ensure system reliability, performance, and uptime by monitoring, analyzing, and automating responses to IT incidents. They integrate data from logs, metrics, and traces to gain real-time insights, helping organizations quickly detect and resolve issues. This role combines expertise in software engineering, monitoring solutions, automation, and machine learning to improve the overall health and efficiency of IT environments.

What are the key skills and qualifications needed to thrive as an Observability AIOps engineer?

To thrive as an Observability AIOps Engineer, you need expertise in systems monitoring, data analytics, automation, and a strong understanding of IT infrastructure, often supported by a degree in computer science or a related field. Familiarity with tools like Prometheus, Grafana, ELK stack, Splunk, and AIOps platforms, as well as certifications in cloud solutions (AWS, Azure, or GCP), are typically required. Strong problem-solving skills, collaboration, and a proactive mindset help you stand out in identifying and addressing system anomalies. These skills and qualities are crucial for maintaining high system reliability, reducing downtime, and enabling data-driven decision-making in complex IT environments.

What are some common challenges faced by Observability AIOps engineers in integrating monitoring solutions across diverse technology stacks?

Observability AIOps Engineers often encounter challenges when integrating monitoring and analytics tools across a mix of legacy systems, cloud-native applications, and various third-party platforms. Ensuring consistent data collection, normalization, and visualization can be complex due to differing protocols, data formats, and tool compatibility. Collaboration with development, operations, and security teams is crucial to address these challenges, streamline workflows, and maintain a unified observability platform. Staying current with evolving AIOps technologies and best practices is also vital for continued success in this dynamic role.

What is the difference between Observability Aiops Engineer vs Site Reliability Engineer?

AspectObservability Aiops EngineerSite Reliability Engineer
Primary FocusMonitoring, analyzing, and improving system observability using AI and automationEnsuring system reliability, scalability, and performance of services
Skills & CertificationsKnowledge of AI/ML, monitoring tools, scripting, cloud platformsSystems engineering, scripting, cloud infrastructure, incident management
Work EnvironmentDevOps teams, monitoring platforms, AI toolsOperations, development teams, cloud environments
Industry UsageTech companies, cloud providers, organizations focusing on AI-driven monitoringLarge-scale tech firms, SaaS providers, internet services

While both roles focus on system performance and reliability, the Observability Aiops Engineer specializes in leveraging AI and automation to enhance system observability, whereas the Site Reliability Engineer concentrates on maintaining overall system stability and scalability. Both roles often collaborate but have distinct core responsibilities.

What are popular job titles related to Observability Aiops Engineer jobs in Ohio?

For Observability Aiops Engineer jobs in Ohio, the most frequently searched job titles are:

What job categories do people searching Observability Aiops Engineer jobs in Ohio look for?

The top searched job categories for Observability Aiops Engineer jobs in Ohio are:

What cities in Ohio are hiring for Observability Aiops Engineer jobs?

Cities in Ohio with the most Observability Aiops Engineer job openings:

Infographic showing various Observability Aiops Engineer job openings in Ohio as of August 2026, with employment types broken down into 87% Full Time, 7% Part Time, and 6% Contract. Highlights an 82% Physical, 5% Hybrid, and 13% Remote job distribution.

Platform Operations Engineer (Site Reliability Engineer)

Westerville, OH • On-site


Vertiv Co

6.9

Company rating: 6.9 out of 10

Based on 66 frontline employees who took The Breakroom Quiz

370th of 495 rated machine equipment manufacturers

Paid breaks

Respectful managers

Uninterrupted breaks


$55 - $73/hr

Full-time

Re-posted 17 days ago


Job description

Job Summary

Vertiv is seeking a skilled Platform Operations Engineer (Site Reliability Engineer) to serve as the owner of cross-platform observability, incident management, and operational reliability within Vertiv's Digital organization. This individual contributor role is responsible for designing, implementing, and continuously improving monitoring and alerting solutions across Vertiv's digital platform ecosystem - including Compass AI, Writer AI, Site Scope, UiPath, Workato, Cursor, and other approved enterprise tools - while owning incident response processes, SLA management, and operational governance. The Platform Operations / SRE will operate within the Digital organization and play a central role in advancing Vertiv's Operational Excellence strategic priority by ensuring the availability, performance, and resilience of platforms that power critical digital workflows and business functions.

As an individual contributor in a lead capacity, this role includes proactive reliability engineering - applying SRE principles such as SLOs, error budgets, and blameless post-mortems - and embedding secure coding and operational governance practices across the Digital organization. The Platform Operations / SRE Engineer will define and enforce observability standards, lead incident response and root cause analysis, manage platform-level SLAs, and partner with engineering, security, and business stakeholders to ensure that all digital platforms meet agreed availability and performance targets. 

This position partners closely with IT Security, NPDI, Digital delivery teams, and business operations, and is based on site at Vertiv's Westerville, OH headquarters.

Responsibilities

  • Own Cross-Platform Monitoring & Observability: Design, implement, and maintain end-to-end monitoring, alerting, and observability solutions across Vertiv's digital platform ecosystem - including AI platforms, automation tools, and internal applications - ensuring real-time visibility into system health, performance, and availability.
  • Lead Incident Response & Management: Serve as the primary escalation point and incident commander for P1/P2 incidents across Digital platforms; lead root cause analysis (RCA), blameless post-mortems, and corrective action tracking to prevent recurrence and reduce mean time to resolution (MTTR).
  • Manage Platform SLAs & Reliability Targets: Define, instrument, and enforce service level objectives (SLOs), service level indicators (SLIs), and error budgets across Digital platforms; produce regular SLA performance reports for leadership and drive platform improvements to meet or exceed agreed availability and performance targets.
  • Drive Secure Coding & Operational Governance: Champion secure coding practices and DevSecOps standards within Digital delivery teams; conduct operational readiness reviews for new platform deployments, enforce configuration management and change control processes, and partner with IT Security and NPDI to ensure all platforms meet Vertiv's security and compliance requirements.
  • Automate Operations & Reduce Toil: Identify and eliminate manual operational toil through automation. This includes automated remediation runbooks and anomaly detection through the use of scripting, IaC tools, and approved automation platforms.
  • Capacity Planning & Performance Engineering: Analyze platform utilization trends and conduct capacity planning across Digital environments; proactively identify performance bottlenecks and recommend architectural improvements to ensure platforms scale reliably with business demand.
  • CI/CD Pipeline Reliability & Deployment Support: Partner with Digital delivery teams to ensure CI/CD pipelines are instrumented for reliability, deployment risk is managed through progressive rollout strategies, and production deployments are supported with appropriate rollback and health-check capabilities.
  • Evaluate & Advance Observability Tooling: Stay current on advancements in observability, AIOps, and SRE tooling; evaluate and recommend new tools and practices that enhance Vertiv's platform operations maturity, and drive adoption of modern reliability engineering standards across the Digital organization.

Requirements

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; equivalent practical experience considered.
  • 5+ years of professional experience in platform operations, site reliability engineering, DevOps, or a related software/infrastructure engineering discipline.
  • 3+ years of hands-on experience with enterprise monitoring and observability platforms (e.g., Datadog, Grafana, Prometheus, Azure Monitor, Splunk, or equivalent) in a multi-platform environment.
  • Demonstrated experience owning and managing incident response processes, post-mortem facilitation, and SLA/SLO governance.
  • Experience implementing secure coding practices, DevSecOps standards, or operational governance frameworks in an enterprise software delivery environment.

Technical Skills

  • Proficiency with monitoring and observability tools (Datadog, Grafana, Prometheus, Azure Monitor, Splunk, or equivalent) for cross-platform health and performance tracking.
  • Strong knowledge of SRE principles, including SLOs, SLIs, blameless post-mortems, and toil reduction practices.
  • Hands-on experience with cloud platforms (AWS preferred) and familiarity with containerized environments (Docker, Kubernetes) and infrastructure-as-code tooling (Terraform, Ansible, or equivalent).
  • Proficiency in at multiple programming languages (Python, Ruby, Powershell, Java, Javascript, C#, etc.) for automation and runbook development.
  • Experience with CI/CD platforms (GitLab, Jenkins, GitHub Actions, Azure DevOps, or equivalent) and deployment reliability practices including progressive rollout, feature flags, and automated health checks.

Preferred Qualifications

  • Google SRE certification, AWS DevOps Professional, Azure certifications, or equivalent SRE/cloud operations certification.
  • Experience with AIOps tooling or AI-assisted anomaly detection and automated remediation capabilities.
  • Familiarity with the Vertiv digital platform ecosystem: Workato, UiPath, Power Automate, Compass AI, Writer AI, or Cursor.
  • Experience applying DevSecOps practices, including SAST/DAST scanning, secrets management, and compliance-as-code in enterprise environments.
  • Experience working in Agile/Scrum delivery environments; familiarity with ITIL incident and change management frameworks.

#LI-RB1

The successful candidate will embrace Vertiv's Core Principles & Behaviors to help execute our Strategic Priorities.

OUR CORE PRINCIPLES: Safety. Integrity. Respect. Teamwork. Inclusion.

OUR STRATEGIC PRIORITIES

High-Performance Culture

Customer Focus

Operational Excellence

Innovation

Financial Strength

VERTIV BEHAVIORS

Own it

Act with urgency

Foster a customer-first mindset

Think big and execute

Lead by example

Drive continuous improvement

Learn and seek out development

Promote transparent & open communication

About Vertiv

Vertiv (NYSE: VRT) brings together hardware, software, analytics and ongoing services to enable its customers' vital applications to run continuously, perform optimally and grow with their business needs. Vertiv solves the most important challenges facing today's data centers, communication networks, and commercial and industrial facilities with a portfolio of power, cooling and IT infrastructure solutions and services that extend from the cloud to the edge of the network. Headquartered in Westerville, Ohio, USA, Vertiv employs around 34,000 people and does business in more than 130 countries. Visit Vertiv.com to learn more.

Work Authorization

No calls or agencies please. Vertiv will only employ those who are legally authorized to work in the United States. This is not a position for which sponsorship will be provided. Individuals with temporary visas such as E, F-1, H-1, H-2, L, B, J, or TN or who need sponsorship for work authorization now or in the future, are not eligible for hire.

Equal Opportunity Employer

Vertiv is an Equal Opportunity/Affirmative Action employer. We promote equal opportunities for all with respect to hiring, terms of employment, mobility, training, compensation, and occupational health, without discrimination as to age, race, color, religion, creed, sex, pregnancy status (including childbirth, breastfeeding, or related medical conditions), marital status, sexual orientation, gender identity / expression (including transgender status or sexual stereotypes), genetic information, citizenship status, national origin, protected veteran status, political affiliation, or disability. If you have a disability and are having difficulty accessing or using this website to apply for a position, you can request help by sending an email to help.join@vertiv.com


What Vertiv employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom