1

Ai Reliability Engineer Jobs in Massachusetts (NOW HIRING)

Senior Site Reliability Engineer

Cambridge, MA ยท On-site

$121K - $218K/yr

Our SRE teams solve reliability, security, and usability at scale for our global fleet while ... AI : Enabling our customers to build, secure, and scale AI apps on the world's most distributed ...

Senior Database Reliability Engineer: Job Type: Full-time Location: Remote Job Summary: Join our ... Our AI tools are designed to complement, not replace, human decision-making. Disclaimer The ...

Senior Database Reliability Engineer

Lynn, MA ยท Remote

$220K - $260K/yr

Senior Database Reliability Engineer: Job Type: Full-time Location: Remote Job Summary: Join our ... Our AI tools are designed to complement, not replace, human decision-making. Disclaimer The ...

Senior Database Reliability Engineer: Job Type: Full-time Location: Remote Job Summary: Join our ... Our AI tools are designed to complement, not replace, human decision-making. Disclaimer The ...

Sr. Site Reliability Engineer

Waltham, MA ยท On-site

$61.50 - $81.75/hr

SS&C is a leading provider of mission-critical, AI-powered technology and services empowering ... Site Reliability Engineer Location(s) : Waltham, MA | Hybrid About the Role Sr Site Reliability ...

Sr. Site Reliability Engineer

Waltham, MA ยท Hybrid

$61.50 - $81.75/hr

SS&C is a leading provider of mission-critical, AI-powered technology and services empowering ... Site Reliability Engineer Location(s) : Waltham, MA | Hybrid About the Role Sr Site Reliability ...

Senior Database Reliability Engineer: Job Type: Full-time Location: Remote Job Summary: Join our ... Our AI tools are designed to complement, not replace, human decision-making. Disclaimer The ...

Showing results 41-60

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What job categories do people searching Ai Reliability Engineer jobs in Massachusetts look for?

The top searched job categories for Ai Reliability Engineer jobs in Massachusetts are:

What cities in Massachusetts are hiring for Ai Reliability Engineer jobs?

Cities in Massachusetts with the most Ai Reliability Engineer job openings:

Remote SRE II for AI Infrastructure & Kubernetes

Akamai Technologies GmbH

Cambridge, MA โ€ข On-site

$125 - $150/hr

Other

Medical, Retirement, PTO

Posted 6 days ago


Key responsibilities

  • Build and maintain dashboards, alerts, and monitoring for inference workloads using Akamai's observability platform

  • Write automation and tooling in Python or Go to reduce operational toil and improve system reliability

  • Participate in on-call rotations, respond to production incidents, and conduct blameless post-mortems


Job description

Are you passionate about cutting-edge AI infrastructure?

Do you want to build your SRE career on one of the most exciting platforms in cloud computing?

Join the Akamai Inference Cloud Team

The Akamai Inference Cloud team is part of Akamai's Cloud Technology Group. We design, implement, deploy and operate AI platforms that enable customers to run inference models and developers to create AI applications.

Partner with the best

In this role, responsibilities will include automation, monitoring, incident response, and working collaboratively with skilled team members. Candidates should possess expertise in Linux systems, automation, and SRE practices. Daily activities involve coding, improving dashboards, enhancing alerts, and minimizing repetitive tasks. Opportunities exist to focus on GPU infrastructure, Kubernetes, and ensuring reliability for AI workloads within Akamai's serverless inference platform.

As an Site Reliability Engineer II, you will be responsible for:

  • Building and maintaining dashboards, alerts, and monitoring for inference workloads using Akamai's existing observability platform
  • Writing automation and tooling in Python or Go to reduce operational toil and improve system reliability
  • Building and improving runbooks for inference-specific operational procedures, integrating into Akamai's existing incident management processes
  • Contributing to SLO tracking and reporting, identifying trends and areas for improvement
  • Supporting CI/CD pipeline maintenance, deployment safety checks, and rollback procedures
  • Collaborating with product engineering teams to troubleshoot complex problems across the stack
  • Participating in on-call rotations, responding to production incidents, and conducting blameless post-mortems

Do what you love

To be successful in this role you will:

  • Have 2+ years of experience in Site Reliability Engineering and a Bachelor's Degree or its equivalent experience
  • Demonstrate coding ability in at least one programming language (Python or Go) with experience writing automation
  • Have experience with Linux systems administration and the ability to troubleshoot complex infrastructure issues
  • Show familiarity with Kubernetes and containerization concepts
  • Have experience with monitoring and observability tools such as Prometheus, Grafana, or similar
  • Have exposure to CI/CD pipelines and infrastructure-as-code tools (Terraform, SaltStack, or equivalent)
  • Show a willingness to learn and grow, with genuine curiosity about AI infrastructure and distributed systems

Work in a way that works for you

FlexBase, Akamai's Global Flexible Working Program, is based on the principles that are helping us create the best workplace in the world. When our colleagues said that flexible working was important to them, we listened. We also know flexible working is important to many of the incredible people considering joining Akamai. FlexBase, gives 95% of employees the choice to work from their home, their office, or both (in the country advertised). This permanent workplace flexibility program is consistent and fair globally, to help us find incredible talent, virtually anywhere. We are happy to discuss working options for this role and encourage you to speak with your recruiter in more detail when you apply.
Learn what makes Akamai a great place to work

Connect with us on social and see what life at Akamai is like!

We power and protect life online, by solving the toughest challenges, together.

At Akamai, we're curious, innovative, collaborative and tenacious. We celebrate diversity of thought and we hold an unwavering belief that we can make a meaningful difference. Our teams use their global perspectives to put customers at the forefront of everything they do, so if you are people-centric, you'll thrive here.

Working for you

At Akamai, we will provide you with opportunities to grow, flourish, and achieve great things. Our benefit options are designed to meet your individual needs for today and in the future. We provide benefits surrounding all aspects of your life:

  • Your health
  • Your finances
  • Your family
  • Your time at work
  • Your time pursuing other endeavors

Our benefit plan options are designed to meet your individual needs and budget, both today and in the future.

About us

Akamai powers and protects life online. Leading companies worldwide choose Akamai to build, deliver, and secure their digital experiences helping billions of people live, work, and play every day. With the world's most distributed compute platform from cloud to edge we make it easy for customers to develop and run applications, while we keep experiences closer to users and threats farther away.

Are you seeking an opportunity to make a real difference in a company with a global reach and exciting services and clients? Come join us and grow with a team of people who will energize and inspire you!
#LI-Remote

Compensation

Akamai is committed to fair and equitable compensation practices. For US based candidates only - the base salary for this position ranges from $95,000 - $171,000/year; a candidateโ€™s salary is determined by various factors including, but not limited to, relevant work experience, skills, certifications and location. Compensation for candidates outside the US will vary. The compensation package may also include incentive compensation opportunities in the form of annual bonus or incentives, equity awards and an Employee Stock Purchase Plan (ESPP). Akamai provides industry-leading benefits including healthcare, 401K savings plan, company holidays, vacation (in the form of PTO), sick time, family friendly benefits including parental leave and an employee assistance program including a focus on mental and financial wellness; Eligibility requirements apply.

Job Info
  • Job Identification 2687
  • Posting Date 03/30/2026, 05:42 PM
  • Locations 145 Broadway, Cambridge, MA, 02142, US (Remote)
#J-18808-Ljbffr