1

Ai Reliability Engineer Jobs (NOW HIRING)

Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

Future Secure AI is building innovative solutions at the forefront of AI technology, seeking a Site Reliability Engineer to design, build, and operate the platforms that power AI Co-Workers. The role ...

Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

Future Secure AI is building innovative solutions at the forefront of AI technology. They are seeking a Site Reliability Engineer to design, build, and operate platforms that support AI Co-Workers ...

NY ยท On-site

$120 - $180/hr

Qualifications * SRE practices, observability (OpenTelemetry, Datadog) * Experience with high-scale distributed systems * Knowledge of AI monitoring (drift, bias, inference metrics) #J-18808-Ljbffr

Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

Future Secure AI is at the forefront of AI technology, tackling significant real-world challenges for global enterprises. They are seeking a Site Reliability Engineer to design, build, and operate ...

About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...

Site Reliability Engineer

New York, NY ยท On-site

$62.25 - $82.75/hr

About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools ... The Role We are seeking highly experienced Site Reliability Engineers (SRE) to shape the ...

Reliability Engineer

Cupertino, CA ยท On-site

$2.0K/mo

About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...

Reliability Engineer

Cupertino, CA ยท On-site

$2.0K/mo

About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...

$110 - $130/hr

Site Reliability Engineer Are you interested in working with the World's leading AI-first Quality Engineering Company? Ready to advance your career, team up with global thought leaders across ...

Digital - Principal SRE (AI Engineer)

Columbus, OH ยท On-site +1

$53.50 - $71.25/hr

The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for ...

Site Reliability Engineer

Austin, TX ยท On-site

$56.50 - $75/hr

About the Role We are looking for a Site Reliability Engineer to help design, build, and operate the platforms that power AI CoWorkers. This is a handson role for an engineer who enjoys owning ...

AI Engineer - Cloud Infrastructure

New York, NY ยท On-site

$175K - $275K/yr

About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise-already trusted by some of the largest companies in the world to troubleshoot, remediate, and even prevent the ...

Digital - Principal SRE (AI Engineer)

Columbus, OH ยท On-site +1

$55 - $73.25/hr

The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for ...

Digital - Principal SRE (AI Engineer)

Columbus, OH ยท On-site +1

$53.50 - $71.25/hr

The Digital - Principal SRE (AI Engineer) role is a position that blends expertise in artificial intelligence, machine learning, and reliability engineering. This professional is responsible for ...

About Traversal Traversal is the AI Site Reliability Engineer (SRE) for the enterprise--already trusted by some of the largest companies in the world to troubleshoot, remediate, and even prevent the ...

Site Reliability Engineer

San Francisco, CA ยท On-site

$150K - $250K/yr

About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to ... As a SRE, you'll be responsible for the reliability, observability, performance, and security of ...

Site Reliability Engineer- AI Enablement

$58.25 - $77.50/hr

Site Reliability Engineer, AI Team: Central AI Location: Remote Travel: None **This position is currently not eligible for visa sponsorship** Be part of a World Class Team responsible for: As a Site ...

Site Reliability Engineer- AI Enablement

$58.25 - $77.50/hr

Site Reliability Engineer, AI Team: Central AI Location: Remote Travel: None **This position is currently not eligible for visa sponsorship** Be part of a World Class Team responsible for: As a Site ...

Showing results 21-40

Ai Reliability Engineer information

See salary details

$61K

$118K

$141K

How much do ai reliability engineer jobs pay per year?

As of Aug 30, 2026, the average yearly pay for ai reliability engineer in the United States is $117,973.00, according to ZipRecruiter salary data. Most workers in this role earn between $102,500.00 and $129,000.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

More about Ai Reliability Engineer jobs

What cities are hiring for Ai Reliability Engineer jobs?

Cities with the most Ai Reliability Engineer job openings:

What states have the most Ai Reliability Engineer jobs?

States with the most job openings for Ai Reliability Engineer jobs include:

Infographic showing various Ai Reliability Engineer job openings in the United States as of August 2026, with employment types broken down into 76% Full Time, 21% Part Time, and 3% Contract. Highlights an 64% Physical, 4% Hybrid, and 32% Remote job distribution, with an average salary of $117,973 per year, or $56.7 per hour.

Site Reliability Engineer

Austin, TX โ€ข On-site

$56.50 - $75/hr

Full-time

Re-posted 22 days ago


Job description

Job Summary:
Future Secure AI is building innovative solutions at the forefront of AI technology, seeking a Site Reliability Engineer to design, build, and operate the platforms that power AI Co-Workers. The role involves owning production infrastructure, enhancing system reliability, and collaborating with engineering teams to ensure robust and scalable systems.
Responsibilities:
โ€ข Design, build, and operate reliable production infrastructure supporting AI Coโ€‘Workers
โ€ข Own Kubernetesโ€‘based platforms used to deploy and run AI workloads
โ€ข Build and maintain infrastructure as code using Terraform
โ€ข Implement and maintain Helmโ€‘based deployment workflows
โ€ข Define, measure, and improve system reliability using SLIs, SLOs, and SLAs
โ€ข Participate in onโ€‘call rotation, incident response, root cause analysis, and postโ€‘mortems
โ€ข Reduce operational toil through automation and engineering improvements
โ€ข Build and improve observability across monitoring, logging, and alerting
โ€ข Partner closely with engineers to ensure systems are resilient, scalable, and secure
โ€ข Operate across build, deploy, and operate phases of the software lifecycle
Qualifications:
Required:
โ€ข Handsโ€‘on Kubernetes experience designing, building, or operating workloads on EKS, AKS, GKE, or selfโ€‘managed Kubernetes
โ€ข Handsโ€‘on Terraform experience for infrastructure provisioning and automation
โ€ข Handsโ€‘on Helm experience for Kubernetes application deployment
โ€ข Professional experience using at least two programming or scripting languages such as Python, Go, Java, Bash, PowerShell, or Ruby
โ€ข Direct Site Reliability Engineer experience or equivalent, including reliability engineering, onโ€‘call, incident response, postโ€‘mortems, and toil reduction
Preferred:
โ€ข Relevant certifications such as CKA, CKAD, cloud certifications, DevOps, DevSecOps, or programming credentials
Company:
Future Secure AI develops secure, bespoke AI Co-Workers and multi-agent systems for complex enterprise workflows. Founded in , the company is headquartered in Austin, USA, with a team of 501-1000 employees. The company is currently Growth Stage.