1

Ai Reliability Engineer Jobs in Massachusetts (NOW HIRING)

Site Reliability Engineer

Cambridge, MA · On-site

$75K - $136K/yr

Our SRE teams solve reliability, security, and usability at scale for our global fleet while ... AI : Enabling our customers to build, secure, and scale AI apps on the world's most distributed ...

Site Reliability Engineer II

Cambridge, MA · On-site

$62.25 - $82.75/hr

Are you passionate about cutting-edge AI infrastructure? Do you want to build your SRE career on one of the most exciting platforms in cloud computing? Join the Akamai Inference Cloud Team The Akamai ...

Site Reliability Engineer

Cambridge, MA · On-site

$75K - $136K/yr

Join our Compute Site Reliability Engineering team Our team is responsible for improving the ... Developing and applying AI-assisted tooling to accelerate incident investigation, identify ...

Do you want to shape reliability practices for a new AI inference platform? Are you a senior ... As a Senior II Site Reliability Engineer, you will be responsible for: * Taking ownership of ...

Join our Compute Site Reliability Engineering team Our team is responsible for improving the ... Developing and applying AI-assisted tooling to accelerate incident investigation, identify ...

New

Lead Site Reliability Engineer

Boston, MA · On-site

$62 - $82.25/hr

At DraftKings, AI is becoming an integral part of both our present and future, powering how work ... The Crown Is Yours As a Lead Site Reliability Engineer, you'll set the reliability standard across ...

Lead Site Reliability Engineer

Boston, MA · On-site

$62 - $82.25/hr

At DraftKings, AI is becoming an integral part of both our present and future, powering how work ... The Crown Is Yours As a Lead Site Reliability Engineer, you'll set the reliability standard across ...

Staff Site Reliability Engineer

Newton, MA · On-site

$62.50 - $83/hr

Manifold is an AI platform for life sciences, focused on accelerating the delivery of life-changing medicines. They are seeking a Staff Site Reliability Engineer to design, build, and operate AWS ...

Join our critical AI Hardware SRE Team! The AI Hardware SRE team is responsible for overseeing, scaling, and optimizing our next-generation dedicated AI hardware infrastructure. You will be ...

We're building revolutionary robotic systems that combine AI, sophisticated control systems, and ... Key job responsibilities As a System Reliability Engineer you will enable development of robotic ...

Showing results 21-40

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What job categories do people searching Ai Reliability Engineer jobs in Massachusetts look for?

The top searched job categories for Ai Reliability Engineer jobs in Massachusetts are:

What cities in Massachusetts are hiring for Ai Reliability Engineer jobs?

Cities in Massachusetts with the most Ai Reliability Engineer job openings:

Senior Site Reliability Engineer

ISO New England Inc.

Holyoke, MA • On-site

$150 - $200/hr

Other

Medical, Dental, Vision, Life, Retirement, PTO

Posted 19 days ago


Job description

ISO New England Inc., One Sullivan Road, Holyoke, Massachusetts, United States of America

Job Description

Posted Monday, August 17, 2026 at 4:00 AM

ISO New England is the independent system operator responsible for ensuring the safe and reliable flow of electricity in our region and planning for the future of the electric grid. We are at the forefront of New England’s ongoing transition to clean energy.

The Senior Site Reliability Engineer (SRE) is a hands-on engineering role responsible for improving the reliability, observability, performance, and operational efficiency of ISO New England's IT services. The SRE works across infrastructure, platform, cyber security, and application teams to reduce operational toil, improve service resilience, and implement scalable automation solutions.

This role has a strong emphasis on observability engineering, automation, Splunk administration, and Infrastructure as Code (IaC). The ideal candidate will possess hands-on experience with Splunk or demonstrate a strong willingness to develop expertise in the platform. Experience with Terraform, automation technologies such as Python and PowerShell, and the ability to leverage AI-assisted development tools to accelerate engineering solutions are key components of the role.

What we offer you:

  • A stable, mission-driven workplace where your impact truly matters
  • A highly engaged work environment that values inclusion, collaboration, and employee safety and wellbeing
  • Competitive compensation with a base salary + performance bonus
  • Robust benefits package, including:
  • Enhanced 401(k) and financial planning support
  • Tuition reimbursement and professional development
  • Wellness programs, including an onsite gym
  • Employee Business Networks
  • Free coffee at our onsite café
  • Hybrid work environment (3 days/week onsite)
  • Distance-based relocation assistance available

How you will make an Impact

  • Build and maintain observability, monitoring, logging, alerting, and telemetry platforms (e.g., Splunk, Dynatrace, PRTG, OpsGenie, StatusPage)
  • Administer, maintain, automate, and continuously improve the Splunk platform, including data onboarding, indexing, search performance, dashboards, access controls, health monitoring, platform scalability, and operational workflows
  • Develop and automate Splunk onboarding, configuration, monitoring, and operational workflows to improve platform reliability and reduce administrative overhead
  • Develop meaningful KPIs and dashboards for business and IT service health
  • Engineer and implement resilience patterns including HA, DR, and automated failover
  • Partner with infrastructure and application teams to plan and execute resilience testing and failover exercises to validate recovery capabilities and observability coverage
  • Conduct performance testing, capacity modeling, forecasting, and right-sizing
  • Participate in major incident response activities, providing technical expertise to accelerate service restoration and identify reliability improvements
  • Identify, prioritize, and eliminate manual operational toil through automation, targeting workflows, runbooks, alerting, platform administration, service management processes, and KPI collection, with a bias toward scalable and repeatable engineering solutions
  • Design, develop, maintain, and support automation solutions, integrations, and operational tooling using Python, PowerShell, Bash, or similar technologies to improve reliability, reduce manual effort, and enhance operational efficiency
  • Design, deploy, and manage infrastructure using Terraform and Infrastructure as Code (IaC) practices, including observability platforms, infrastructure services, and supporting technology stacks, with a focus on consistency, repeatability, and operational sustainability
  • Identify gaps in observability coverage and drive engineering solutions to close them
  • Collaborate with architecture and application teams to ensure production readiness
  • Leverage AI-assisted development tools to accelerate automation initiatives while reviewing, validating, troubleshooting, and refining generated code to ensure reliability, security, maintainability, and operational effectiveness
  • Reduce repeat incidents by engineering permanent fixes and driving continuous improvement

What we are looking for

  • 5+ years of experience in SRE, DevOps, systems engineering, platform engineering, or IT operations
  • Experience with enterprise monitoring and observability platforms. Hands-on experience with Splunk is strongly preferred. Candidates without direct Splunk experience must demonstrate a strong willingness and aptitude to develop expertise in Splunk administration, engineering, and automation.
  • Experience designing, deploying, or managing infrastructure using Terraform and Infrastructure as Code (IaC) practices
  • Strong scripting and automation experience using Python, PowerShell, Bash, or similar technologies, including the development of operational tooling, integrations, and workflow automation in production environments
  • Demonstrated experience designing, developing, and supporting automation solutions that measurably reduced manual operational effort in an enterprise environment
  • Ability to read, understand, review, troubleshoot, and refine code produced by engineering teams or AI-assisted development platforms
  • Knowledge of distributed systems, networking, enterprise infrastructure, and cloud platforms
  • Familiarity with SRE principles including SLOs, error budgets, observability, and toil reduction
  • Ability to analyze and troubleshoot complex technical systems
  • Preferred Qualifications
  • Experience in mission-critical, highly available, or regulated environments
  • Experience utilizing AI-assisted development tools to accelerate automation, operational engineering, or platform management activities
  • Knowledge of ITIL processes and/or SRE best practices
  • Experience with performance testing, capacity planning, resilience testing, or disaster recovery validation

This employer will not sponsor applicants for work visas for this position (ex: H-1B, F-1/CPT/OPT, O-1, E-3, TN, J, etc.).

The expected salary range for this position is $134,000 - $170,000 per year, for a Senior to Lead level candidate. This role is also eligible for an annual performance bonus, comprehensive health insurance (medical, dental and vision), flexible spending and health savings accounts, a 401(k) plan with generous employer contributions and a student debt benefit, life and AD&D insurance, disability insurance, critical illness and hospital indemnity benefits, paid time off, paid leave, a wellness program, an employee assistance program and other great company perks.

#LI-HYBRID

This is a U.S. based role. If the successful candidate resides outside of the U.S., relocation will be required.

Equal Opportunity: We are proud to be an EEO employer. Applicants for employment are considered without regard to race, color, religion, creed, sex (including pregnancy, childbirth, and related medical conditions), gender identity or expression, sexual orientation, citizenship, national origin, age, ancestry, marital status, disability (including learning, mental, intellectual, and physical), service in the uniformed services, genetic information, or any other status protected by applicable law.

Drug Free Environment: We maintain a drug-free workplace and perform pre-employment substance abuse testing.

ISO New England Inc., One Sullivan Road, Holyoke, Massachusetts, United States of America

#J-18808-Ljbffr