The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and ...
The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and ...
Staff Software Engineer, AI Reliability
$67.25 - $89.25/hr
About the Role AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API ...
Staff Software Engineer, AI Reliability
$67.25 - $89.25/hr
About the Role AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API ...
Member of Technical Staff, AI Reliability & Monitoring Engineering Lead
San Francisco, CA · On-site
$67.25 - $89.25/hr
They are seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of ...
Member of Technical Staff, AI Reliability & Monitoring Engineering Lead
San Francisco, CA · On-site
$67.25 - $89.25/hr
They are seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of ...
Staff Software Engineer, AI Reliability
San Francisco, CA · On-site
$325K - $485K/yr
About the Role AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API ...
Staff Software Engineer, AI Reliability
San Francisco, CA · On-site
$325K - $485K/yr
About the Role AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API ...
Senior Platform & Reliability Engineer
San Francisco, CA · On-site
$300K - $400K/yr
Senior Platform & Reliability Engineer About OpenArt OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We're building the next generation of creative tools ...
Senior Platform & Reliability Engineer
San Francisco, CA · On-site
$300K - $400K/yr
Senior Platform & Reliability Engineer About OpenArt OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We're building the next generation of creative tools ...
Founding Platform & Reliability Engineer
San Francisco, CA · On-site
$67.25 - $89.25/hr
Founding Platform & Reliability Engineer About OpenArt OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We're building the next generation of creative tools ...
Founding Platform & Reliability Engineer
San Francisco, CA · On-site
$67.25 - $89.25/hr
Founding Platform & Reliability Engineer About OpenArt OpenArt is an AI Storytelling and Visual Creation Platform used by millions worldwide. We're building the next generation of creative tools ...
About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...
About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...
About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...
About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...
Reliability Engineer
Cupertino, CA · On-site
$2.0K/mo
About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...
Reliability Engineer
Cupertino, CA · On-site
$2.0K/mo
About Etched Etched is building AI chips that are hard-coded for individual model architectures ... Reliability Engineer We are seeking a skilled and detail-oriented Reliability Engineer to join our ...
Head of SRE
Palo Alto, CA · On-site
$67 - $89.25/hr
Wand AI is a company focused on integrating AI into the workforce, enabling humans and AI agents to work together efficiently. They are seeking a hands-on Head of SRE to establish and lead their Site ...
Head of SRE
Palo Alto, CA · On-site
$67 - $89.25/hr
Wand AI is a company focused on integrating AI into the workforce, enabling humans and AI agents to work together efficiently. They are seeking a hands-on Head of SRE to establish and lead their Site ...
Site Reliability Engineer
San Francisco, CA · On-site
$67.25 - $89.25/hr
About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to ... As a SRE, you'll be responsible for the reliability, observability, performance, and security of ...
Site Reliability Engineer
San Francisco, CA · On-site
$67.25 - $89.25/hr
About Runloop Runloop.ai is pioneering the next generation of infrastructure and orchestration to ... As a SRE, you'll be responsible for the reliability, observability, performance, and security of ...
Member of Technical Staff, AI Reliability & Monitoring Engineering Lead
San Francisco, CA · On-site
$256K - $276K/yr
The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and ...
Member of Technical Staff, AI Reliability & Monitoring Engineering Lead
San Francisco, CA · On-site
$256K - $276K/yr
The Opportunity Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and ...
Reliability Engineer
$70 - $115/hr
Description Reliability Engineering team works across client's entire product portfolio and ... AI tools.
New
Quick apply
Reliability Engineer
$70 - $115/hr
Description Reliability Engineering team works across client's entire product portfolio and ... AI tools.
New
We believe AI will fundamentally reshape how SRE is practiced - from incident detection and resolution to capacity planning and toil elimination - and we're looking for a leader who shares that ...
We believe AI will fundamentally reshape how SRE is practiced - from incident detection and resolution to capacity planning and toil elimination - and we're looking for a leader who shares that ...
Site Reliability Engineer - AI Infrastructure
San Francisco, CA · On-site +1
$67.25 - $89.25/hr
Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco • Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage ...
Site Reliability Engineer - AI Infrastructure
San Francisco, CA · On-site +1
$67.25 - $89.25/hr
Site Reliability Engineer - AI Infrastructure Location: Global Remote / San Francisco • Full-Time About Andromeda Andromeda Cluster was founded by Nat Friedman and Daniel Gross to give early-stage ...
Site Reliability Engineer (SRE)
San Francisco, CA · On-site
$116K - $200K/yr
We're a family-founded company on a mission to create the world's first AI-powered Personal ... The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the ...
Site Reliability Engineer (SRE)
San Francisco, CA · On-site
$116K - $200K/yr
We're a family-founded company on a mission to create the world's first AI-powered Personal ... The Role As a Site Reliability Engineer (SRE) at Air Apps, you will be responsible for ensuring the ...
Reliability Engineer
Costa Mesa, CA · On-site
$108K - $136K/yr
Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...
Reliability Engineer
Costa Mesa, CA · On-site
$108K - $136K/yr
Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...
Reliability Engineer
Costa Mesa, CA · On-site
$110K - $138K/yr
Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...
Reliability Engineer
Costa Mesa, CA · On-site
$110K - $138K/yr
Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, CA · On-site
$70.25 - $93.50/hr
We believe AI will fundamentally reshape how SRE is practiced - from incident detection and resolution to capacity planning and toil elimination - and we're looking for a leader who shares that ...
Site Reliability Engineering (SRE) Manager, Apple Maps
Cupertino, CA · On-site
$70.25 - $93.50/hr
We believe AI will fundamentally reshape how SRE is practiced - from incident detection and resolution to capacity planning and toil elimination - and we're looking for a leader who shares that ...
Site Reliability Engineer
Mountain View, CA · Hybrid
$189K - $232K/yr
As a Site Reliability Engineer, you will strengthen infrastructure, optimize tooling, deepen ... You will harness AI-assisted development and operational workflows to minimize toil, accelerate ...
Site Reliability Engineer
Mountain View, CA · Hybrid
$189K - $232K/yr
As a Site Reliability Engineer, you will strengthen infrastructure, optimize tooling, deepen ... You will harness AI-assisted development and operational workflows to minimize toil, accelerate ...
Ai Reliability Engineer information
What are the key skills and qualifications needed to thrive as an AI Reliability Engineer, and why are they important?
What is the difference between Ai Reliability Engineer vs Data Scientist?
| Aspect | Ai Reliability Engineer | Data Scientist |
|---|---|---|
| Required Credentials | Bachelor's or master's in CS, engineering, or related; certifications in AI/ML | Bachelor's or master's in CS, statistics, or related; certifications in data analysis or ML |
| Work Environment | Tech companies, AI-focused teams, engineering departments | Research labs, tech firms, analytics teams |
| Employer & Industry Usage | AI product development, machine learning systems, reliability testing | Data analysis, predictive modeling, business insights |
While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.
What are AI Reliability Engineers?
What are some common challenges Ai Reliability Engineers face when ensuring model robustness in production environments?
- Senior Remote Net Core Developer
- Entry Level Devops Engineer
- Overnight Site Reliability Engineer Remote
- Temp Junior Site Reliability Engineer
- Trainee Devops Engineer
- Senior Reliability Engineer
- Ai Devops Engineer
- Work From Home Salesforce Marketing Cloud Developer
- Silicon Foundry Interface Engineer
- Entry Level Ai Infrastructure Engineer

$256K - $276K/yr
Other
Posted 19 days ago
Job description
Postman is seeking an experienced AI Systems Reliability Engineer to help define, build, and maintain the infrastructure and processes that ensure the reliability, scalability, and performance of Postman's AI-powered API and agentic systems in production. This role focuses on monitoring, availability, incident response, and automation to support AI services and tools trusted by millions of developers globally.
What You'll DoDevelop and manage reliability metrics (SLOs) for AI-driven API services and agentic AI platform features
Implement comprehensive observability and monitoring systems for real-time performance and fault detection
Design and drive automated failover, recovery, and incident response strategies for high-availability AI infrastructure
Optimize resource utilization, particularly GPU/accelerator efficiency, ensuring cost-effective AI system operation
Collaborate closely with engineering, platform, and product teams to align reliability efforts with broader organizational goals
Lead efforts to build internal tooling and automation focused on AI system stability and operational excellence
Drive continuous improvement in deployment practices, monitoring approaches, and incident management processes
Have a strong background in AI reliability engineering, SRE, or DevOps for distributed systems
Understand the unique challenges of maintaining large-scale AI systems and integrating AI-specific metrics into reliability frameworks
Are experienced with cloud platforms, monitoring tools, and incident response automation
Are comfortable collaborating across teams to influence best practices for AI system reliability and operational health
Thrive in dynamic, fast-paced environments focusing on delivering reliable, safe AI-powered services
Bonus Skills and Experiences
Hands-on experience with AI/ML infrastructure, including GPU/xPU optimization and scaling
Familiarity with API platform operations and large-scale distributed services
Prior experience building or operating observability tools tailored for AI and agentic systems
Contribution to open-source projects or reliability engineering thought leadership
The reasonably estimated base salary for this role ranges from $256,000 to $276,000, plus a competitive equity package. Actual compensation is based on the candidate's skills, qualifications, and experience.
About Postman
Sourced by ZipRecruiter
Industry
Software development
Company size
501 - 1,000 Employees
Headquarters location
San Francisco, CA, US
Year founded
2014