1

Ai Reliability Engineer Jobs in Portland, OR (NOW HIRING)

This role emphasizes reliability, control, observability, data quality, and governance for agentic ... DevOps and Reliability for AI Agent Systems Define and track SLIs/SLOs for task completion ...

We believe transformative AI should have a positive impact on people-powerful in capability, yet ... Shipping and running it reliably in production is owned by the SRE / Production squad - you partner ...

Sr. AI Software Engineer - Agent Harness

Hillsboro, OR ยท On-site

$133K - $175K/yr

We believe transformative AI should have a positive impact on people-powerful in capability, yet ... Shipping and running it reliably in production is owned by the SRE / Production squad - you partner ...

Software Engineer-NetENG

Portland, OR ยท On-site

$114K - $142K/yr

At New Relic, we build an intelligent platform that empowers companies to thrive in an AI-first ... reliability that the entire New Relic organization depends on to deliver real-time insights. You ...

The AI Legal Engineer is responsible for partnering with subject matter experts to design, build ... Build and execute evaluation frameworks to measure workflow quality, accuracy, and reliability ...

Senior Applied AI Engineer

Hillsboro, OR ยท On-site

$113K - $156K/yr

Drive improvements in architecture, performance, and reliability, enabling teams to bring to bear ... MS or higher degree (or equivalent experience) in Computer Science, Engineering, AI, or a related ...

Associate IT DevOps Engineer

Beaverton, OR ยท Remote

$80K - $95K/yr

Enabled by modern platforms and AI, you'll do the most meaningful work of your career and see your ... Contribute to platform architecture, reliability, and lifecycle management Infrastructure as Code ...

Lead AI/ML Solutions Engineer

Gresham, OR ยท On-site

$108K - $143K/yr

... ensuring security, reliability, maintainability, and responsible AI practices. Main ... Collaborate with platform engineering teams to leverage approved enterprise AI technologies and ...

next page

Showing results 1-20

Ai Reliability Engineer information

See Portland, OR salary details

$64.7K

$125.1K

$149.5K

How much do ai reliability engineer jobs pay per year?

As of Jul 28, 2026, the average yearly pay for ai reliability engineer in Portland, OR is $125,111.00, according to ZipRecruiter salary data. Most workers in this role earn between $108,700.00 and $136,800.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as an AI Reliability Engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are AI Reliability Engineers?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges Ai Reliability Engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What are popular job titles related to Ai Reliability Engineer jobs in Portland, OR? For Ai Reliability Engineer jobs in Portland, OR, the most frequently searched job titles are:
What job categories do people searching Ai Reliability Engineer jobs in Portland, OR look for? The top searched job categories for Ai Reliability Engineer jobs in Portland, OR are:
Senior Site Reliability Engineer (REMOTE)

Senior Site Reliability Engineer (REMOTE)

Discogs

Beaverton, OR โ€ข Remote

$140K/yr

Full-time

Medical, Dental, Retirement, PTO

Posted 20 days ago


Job description

Senior Site Reliability Engineer (REMOTE)

Department: Engineering

Employment Type: Full Time

Location: REMOTE US

Compensation: $140,000 / year


Description Who Weโ€™re Looking For
The Discogs Platform team is focused on several objectives: building and supporting performant, cost-effective, reliable cloud infrastructure; data administration of complex CDC workflows; and developer experience tooling and mentorship, including our agentic AI discipline. As a Platform member, the Senior Site Reliability Engineer will contribute to the Platform teamโ€™s centralized infrastructure, including maintenance, monitoring, and automation of services ranging from databases to Kubernetes; lead incident response and postmortem efforts; and work closely with other engineering teams to understand their needs and drive improvements to both our technologies and processes.
Location While we are a remote company we are only hiring for the following locations: OR, WA, CA, CO, TX, IL
Compensation Fixed Position Rate: $140,000
This role carries a single, non-negotiable Fixed Position Rate to ensure absolute equity and eliminate negotiation bias.
Key ResponsibilitiesWhat Youโ€™ll AccomplishReasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
  • Owning tasks and larger projects from planning to production rollout
  • Learning new technologies and building expertise with the goal of teaching and mentoring others; mentoring with the goal of force-multiplying through docs and tools
  • Maintaining organization cloud presence in AWS
  • Automating and deploying infrastructure configurations using Infrastructure as Code (IAC)
  • Mentoring engineering squads on Platform best practices for Kubernetes, MySQL, Kafka, and other software development lifecycle areas
  • Assisting engineering squads with capacity planning, on-call preparation, and production readiness
  • Writing documentation and runbooks that contribute to the engineering organizationโ€™s knowledge base
  • Implementing monitoring and alerting systems with Discogs observability tools
  • Working in a containerized, orchestrated environment
  • Participating in on-call rotation, responding to incidents, and troubleshooting data and other operations issues
  • Contributing to the reliability and design patterns of our Kafka CDC and event workflows
  • Contributing to agentic AI best practices and tooling, including skills, agents, and safety

Skills, Knowledge and ExpertiseWhat Youโ€™ll Contribute
Required (or equivalent tools; listed is our current stack):
  • Infrastructure-as-code (Terraform)
  • CI/CD (GitHub Actions)
  • Kubernetes (EKS, Kustomize, Karpenter, administration, application manifests)
  • AWS and cloud development (VPC, EKS, RDS, S3)
  • FinOps and cloud cost optimization
  • Observability (Datadog, Sentry)
  • Agentic AI (Claude Code)
  • Scripting (Shell, Python)
  • Track record of collaboration and mentorship
  • Excellent written communication and documentation skills
  • Continuous learning
  • Ownership and proactive approach to solving large problems
Preferred:
  • Kafka: Cluster administration (Strimzi), Kafka Connect (Debezium, JDBC)
  • Flink
  • Relational database administration and performance (MySQL, Percona Server, AWS RDS)
  • Elasticsearch (ECK administration, scaling, performance)
  • Python (SQLAlchemy, FastAPI)
  • GraphQL (schema design, Apollo federation)
  • REST API
  • GitOps (ArgoCD)
  • Hashicorp Vault
  • Redis
  • Memcached
Education & Experience:
  • A Bachelor\'s Degree in Computer Science or similar area of focus, or equivalent relevant work experience.
  • 5+ years experience in Ops, DevOps, Site Reliability, Platform or other systems roles.

BenefitsWhat We Provide
  • Competitive compensation: salary, plus performance-related bonus program
  • 401(k) with employer match
  • 100% company-paid medical and dental insurance benefits for you and your dependents
  • 4 weeks paid vacation, increasing based on tenure
  • 18 weeks paid leave for birth moms
  • 8 weeks paid parental leave, including for adoption
  • Monthly wellness allowance
  • Annual professional and personal development allowance
  • Work from home office set-up and expense allowances 
  • Flexible work location opportunities 
  • Employer matching toward charitable contributions