1

Ai Reliability Engineer Jobs in New York (NOW HIRING)

Site Reliability Engineer

New York, NY · On-site

$62.25 - $82.75/hr

  • Medical

  • Retirement

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools ... The Role We are seeking highly experienced Site Reliability Engineers (SRE) to shape the ...

AI Engineer - Cloud Infrastructure

New York, NY · On-site

$175K - $275K/yr

  • Medical

About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise-already trusted by some of the largest companies in the world to troubleshoot, remediate, and even prevent the ...

Site Reliability Engineer (SRE)

Parsippany, NJ · On-site

$57.25 - $76.25/hr

We are looking for a talented Site Reliability Engineer (SRE) with a strong background in Google ... Familiarity with Google BI and AI/ML tools a plus (Looker, BigQuery ML, Vertex AI, etc.) Experience ...

Site Reliability Engineer

Jersey City, NJ · On-site

$59.50 - $79/hr

  • Medical

  • Dental

  • Vision

Our AI platform, 1Exiger, delivers instant visibility into complex supplier ecosystems, leveraging ... Site Reliability Engineer Location: U.S. (Hybrid) This role requires U.S. citizenship and ...

Site Reliability Engineer III

Jersey City, NJ · On-site

$62.25 - $82.75/hr

  • Medical

  • Retirement

As a Site Reliability Engineer III at JPMorgan Chase within the within the Consumer & Community ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Site Reliability Engineer III

Jersey City, NJ · On-site

$62.25 - $82.75/hr

  • Medical

  • Retirement

As a Site Reliability Engineer III at JPMorgan Chase within the within the Consumer & Community ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

GCP Site Reliability Engineer Interview Mode: candidates local to Parsippany, Nj who can attend an ... Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc.) Experience with ...

Site Reliability Engineer III

Jersey City, NJ · On-site

$62.25 - $82.75/hr

  • Medical

  • Retirement

As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Site Reliability Engineer III

Jersey City, NJ · On-site

$59.50 - $79/hr

  • Medical

  • Retirement

As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Staff Site Reliability Engineer (SRE) (Hybrid)

Holmdel, NJ · Hybrid

$63.75 - $84.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

The Splunk Agent Resilience team is defining the future of AI resilience. Together our team ... As a Staff Site Reliability Engineer (SRE), you will provide technical leadership for the ...

Site Reliability Engineer III

Jersey City, NJ · On-site

$62.25 - $82.75/hr

  • Medical

  • Retirement

As a Site Reliability Engineer III at JPMorgan Chase within the Chief Data & Analytics Office (CDAO ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Mid-Level Site Reliability Engineer

New York, NY · On-site

$150K - $190K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

  • PTO

Role Overview Instalily, a cutting-edge AI startup, is seeking a curious and highly skilled Site Reliability Engineer to help build the Internal Developer Platform (IDP) that powers our AI agent ...

Staff Site Reliability Engineer (SRE) (Hybrid)

New York, NY · Hybrid

$62.25 - $82.75/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

The Splunk Agent Resilience team is defining the future of AI resilience. Together our team ... As a Staff Site Reliability Engineer (SRE), you will provide technical leadership for the ...

next page

Showing results 1-20

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are popular job titles related to Ai Reliability Engineer jobs in New York?

For Ai Reliability Engineer jobs in New York, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in New York look for?

The top searched job categories for Ai Reliability Engineer jobs in New York are:

What cities in New York are hiring for Ai Reliability Engineer jobs?

Cities in New York with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in New York as of August 2026, with employment types broken down into 70% Full Time, and 30% Contract. Highlights an 90% In-person, and 10% Remote job distribution.

Staff Software Engineer (AI Reliability)

Menlo Ventures

Manhattan, NY • On-site

$325 - $485/hr

Other

Posted 10 days ago


Job description

About The Role

AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths – every hop from the SDK through our network, API layers, serving infrastructure, and accelerators and back. We jump into the trenches alongside partner teams to make the systems that deliver Claude more robust and resilient, be it during an incident or collaborating on projects. Reliability here is an emergent phenomenon that transcends any single team’s boundaries, so someone has to zoom out and look at the whole picture. That’s us – and it means few teams at Anthropic offer this kind of dynamic, cross‑cutting exposure to the systems that matter most.

Responsibilities
  • Develop appropriate Service Level Objectives for large language model serving systems, balancing availability and latency with development velocity
  • Design and implement monitoring and observability systems across the token path
  • Assist in the design and implementation of high‑availability serving infrastructure across multiple regions and cloud providers
  • Lead incident response for critical AI services, ensuring rapid recovery, thorough incident reviews, and systematic improvements
  • Support the reliability of safeguard model serving – critical for both site reliability and Anthropic’s safety commitments
You May Be a Good Fit If You
  • Have strong distributed systems, infrastructure, or reliability backgrounds – we’re looking for reliability‑minded software engineers and SREs
  • Are curious and brave – comfortable jumping into unfamiliar systems during an incident and helping drive resolution even when you don’t have deep expertise yet
  • Think holistically about how systems compose and where the seams are
  • Can build lasting relationships across teams – our engagement model depends on being welcomed as teammates, not outsiders with opinions
  • Care about users and feel ownership over outcomes, even for systems you don’t own
  • Have excellent communication and collaboration skills – you’ll be partnering across the entire company
  • Bring diverse experience – the team’s strength comes from people who’ve built product stacks, scaled databases, run massive distributed systems, and everything in between
Strong Candidates May Also
  • Have been an SRE, Production Engineer, or in similar reliability‑focused roles on large scale systems
  • Have experience operating large‑scale model serving or training infrastructure (over 1000 GPUs)
  • Have experience with one or more ML hardware accelerators (GPUs, TPUs, Trainium)
  • Understand ML‑specific networking optimizations like RDMA and InfiniBand
  • Have expertise in AI‑specific observability tools and frameworks
  • Have experience with chaos engineering and systematic resilience testing
  • Have contributed to open‑source infrastructure or ML tooling
Annual Salary

$325,000—$485,000 USD

Logistics

Minimum education: Bachelor’s degree or an equivalent combination of education, training, and/or experience.

Required field of study: A field relevant to the role as demonstrated through coursework, training, or professional experience.

Minimum years of experience: Years of experience required will correlate with the internal job level requirements for the position.

Location‑based hybrid policy: Currently, we expect all staff to be in one of our offices at least 25% of the time. However, some roles may require more time in our offices.

Visa sponsorship: We do sponsor visas! However, we aren’t able to successfully sponsor visas for every role and every candidate. If we make you an offer, we will make every reasonable effort to get you a visa and retain an immigration lawyer to help with this.

#J-18808-Ljbffr