1

On Call Aiops Engineer Jobs (NOW HIRING)

DevOps Engineer III

Waltham, MA · Hybrid

$57 - $78/hr

Participate in on-call rotations, incident triage, and post-incident reviews; apply SRE best ... Demonstrated experience using AI-powered copilots, chat assistants, or AIOps platforms to ...

Develop automation and AIOps capabilities to reduce toil, accelerate incident detection, and enable ... Track record of incident management and on-call ownership; comfort with incident response and the ...

Senior Systems Engineer

Topeka, KS · On-site

$98K - $134K/yr

... AIOps initiatives. * Cross-team initiatives : Lead initiatives that span Platform Engineering ... Comfortable serving as senior escalation in an on-call rotation. Preferred Qualificatons * Deep ...

Senior Systems Engineer

Atlanta, GA · On-site

$100K - $137K/yr

... AIOps initiatives. * Cross-team initiatives : Lead initiatives that span Platform Engineering ... Comfortable serving as senior escalation in an on-call rotation. Preferred Qualificatons * Deep ...

AWS DevOps Datadog

Sandy Springs, GA · On-site

$52.25 - $71.75/hr

... SRE Engineer - Datadog & AIOps Location: Atlanta, Georgia -- Preferred Onsite/Hybrid Employment ... on-call operations, reliability engineering, capacity planning, and blameless post-incident ...

Senior Systems Engineer

Overland Park, KS · On-site

$103K - $141K/yr

... AIOps initiatives. * Cross-team initiatives : Lead initiatives that span Platform Engineering ... Comfortable serving as senior escalation in an on-call rotation. Preferred Qualificatons * Deep ...

Define and implement extensibility patterns including AIOps: e.g anomaly detection hooks, event ... building, on-call rotation, etc. * Travel maybe required to team or project events. Required ...

Define and implement extensibility patterns including AIOps: e.g anomaly detection hooks, event ... building, on-call rotation, etc. * Travel maybe required to team or project events. Required ...

... AIOps-assisted diagnostics. Architect and support cloud networking across Azure (VNets, VNet ... on-call rotation for critical network incidents. Qualifications 710 years of enterprise network ...

next page

Showing results 1-20

On Call Aiops Engineer information

What are the most commonly searched types of Aiops Engineer jobs?

The most popular types of Aiops Engineer jobs are:

Senior Site Reliability Engineer, AIOPs

NVIDIA

Santa Clara, CA • On-site

$67 - $89/hr

Full-time

Re-posted 29 days ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 18 frontline employees who took The Breakroom Quiz

7th of 246 rated software companies


Job description

Job Summary:
NVIDIA has been transforming computer graphics and computing for over 25 years, and they are seeking a Senior Site Reliability Engineer to join their innovative team. The role involves operating an AI Data Center AIOps platform, ensuring uptime, performance, and data integrity while collaborating with engineering teams to create actionable insights and automation.
Responsibilities:
• Continuously monitor platform health via dashboards/logs/metrics, automate recurring checks, and keep reliability + resource efficiency on track.
• Own Kubernetes deployments end-to-end (runbooks, canary checks, post-deploy validation), and lead rollbacks/remediations when needed.
• Lead first-level incident triage: collect diagnostics, identify likely root causes, and hand off clear, actionable findings to engineering.
• Build and maintain runbooks/SOPs/checklists, pushing continuous improvement through automation.
• Manage deployment infrastructure and packaging (Helm + Terraform/IaC) to keep environments scalable, consistent, and reproducible.
• Contribute in adjacent functional areas to grow and help your team members!
Qualifications:
Required:
• BS/MS in CS/CE (or equivalent experience) and 5+ years operating production distributed systems as SRE/DevOps/Platform Ops.
• Proven ownership of reliability for an observability/AIOps platform: SLOs/SLIs, on-call, addressing incidents, and follow-up evaluations that drive measurable improvements.
• Deep Kubernetes + containers experience (deploying, debugging, scaling) for telemetry-heavy microservices—ingestion, processing, storage, APIs, and UI.
• Automation-first approach: solid scripting (Python/Bash), CI/CD, and infrastructure-as-code (Terraform + Helm) to deliver safe rollouts (canaries/rollbacks), reproducible environments, and minimal toil.
• Clear communicator who writes excellent runbooks/docs and can translate ambiguous requirements into concrete operational practices and dependable customer-facing reliability.
Preferred:
• Strong Linux + networking fundamentals, distributed systems instincts, and hands-on ops for Kubernetes/services/streaming stacks are ideal; bonus for experience with observability platforms at scale.
• Experience building safe automation that operators trust: canary releases, automated rollback criteria, 'monitoring for the monitoring' (lag/drop/error budgets), and replay/backfill pipelines with correctness checks.
• Strong in distributed/streaming systems operations (Kafka/Pulsar, Flink/Spark, ClickHouse/Elastic/TSDBs, object storage)—and can reason about backpressure, hotspots, and failure domains end-to-end.
• Proven programming experience building automation tools or services — ideally in Python, or similar languages — to simplify operations and scale recurring processes.
• Proven experience running large‑scale production deployments and multiple Kubernetes environments or clusters across teams or customers, coordinating changes and rollouts with minimal disruption with hands‑on experience with observability tools — you know your way around dashboards, metrics, logs, and traces using platforms like Prometheus, Grafana, or similar.
Company:
NVIDIA is a computing platform company operating at the intersection of graphics, HPC, and AI. Founded in 1993, the company is headquartered in Santa Clara, USA, with a team of 10001+ employees. The company is currently Late Stage.

What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US

Year founded

1993