1

Ai Reliability Engineer Jobs in Houston, TX (NOW HIRING)

Site Reliability Engineer III

Houston, TX

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Risk ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Site Reliability Engineer III

Houston, TX

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Site Reliability Engineer III

Houston, TX ยท On-site

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Site Reliability Engineer III

Houston, TX ยท On-site

$54.50 - $72.25/hr

As a Site Reliability Engineer III at JPMorgan Chase within the Corporate Technology, Corporate ... Uses enterprise-authorized AI capabilities within the work environment to accelerate incident ...

Reliability Engineer

Houston, TX ยท On-site

$97K - $123K/yr

ON.energy is building the backbone of energy and AI infrastructure powering grid-safe data centers ... Position Summary ON.energy is seeking a Reliability & Procedures Engineer to strengthen how our ...

Reliability Engineer HROC

Houston, TX ยท On-site

$107K - $143K/yr

What You'll Be Doing * Lead and apply regional reliability engineering strategies to improve ... use AI-assisted tools to support the screening and evaluation of candidate applications and ...

Reliability Engineer HROC

Houston, TX ยท On-site

$107K - $143K/yr

What You'll Be Doing * Lead and apply regional reliability engineering strategies to improve ... use AI-assisted tools to support the screening and evaluation of candidate applications and ...

Site Reliability Engineer II

Houston, TX ยท Hybrid

$54.50 - $72.25/hr

Site Reliability Engineer II About PROS: PROS, Inc. is the leading offer management provider to the ... Powered by AI, the PROS Platform enables commercial teams to align capacity with demand and ...

Site Reliability Engineer II

Houston, TX ยท On-site

$54.50 - $72.25/hr

Site Reliability Engineer II About PROS: PROS, Inc. is the leading offer management provider to the ... Powered by AI, the PROS Platform enables commercial teams to align capacity with demand and ...

next page

Showing results 1-20

Ai Reliability Engineer information

See Houston, TX salary details

$58.3K

$112.7K

$134.7K

How much do ai reliability engineer jobs pay per year?

As of Aug 21, 2026, the average yearly pay for ai reliability engineer in Houston, TX is $112,661.00, according to ZipRecruiter salary data. Most workers in this role earn between $97,900.00 and $123,200.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Houston, TX?

For Ai Reliability Engineer jobs in Houston, TX, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Houston, TX look for?

The top searched job categories for Ai Reliability Engineer jobs in Houston, TX are:

What cities near Houston, TX are hiring for Ai Reliability Engineer jobs?

Cities near Houston, TX with the most Ai Reliability Engineer job openings:

Staff Observability Platform Engineer (SRE)

Recruitment.ai

Houston, TX โ€ข On-site

$54.50 - $72.25/hr

Other

Posted 3 days ago

New


Job description

Staff Observability Platform Engineer (SRE)
Locations: Seattle, WA (Hybrid), Houston, TX (Hybrid), New York, NY (Hybrid)
What Youโ€™ll Do
  • Design, build, and evolve observability platforms across metrics, logs, traces, alerting, and telemetry pipelines.
  • Lead the implementation of scalable observability solutions that support Nscaleโ€™s growing GPU and AI infrastructure.
  • Partner with SRE, infrastructure, platform, and AI/ML teams to ensure observability is embedded throughout the software and infrastructure lifecycle.
  • Drive improvements in monitoring coverage, alert quality, service health visibility, and incident response effectiveness.
  • Develop standards, frameworks, and reusable patterns that simplify observability adoption across engineering teams.
  • Identify reliability risks and operational blind spots, helping teams proactively address them before they impact customers.
  • Contribute to architectural decisions around telemetry collection, storage, retention, cardinality management, and performance optimization.
  • Lead technical initiatives and projects that improve platform scalability, reliability, and operational efficiency.
  • Mentor engineers and provide technical guidance through design reviews, code reviews, and knowledge sharing.
  • Participate in incident investigations and postmortems, translating operational learnings into durable platform improvements.
  • Evaluate new observability technologies and practices, balancing innovation with operational simplicity and long-term maintainability.
About You
  • Experience in SRE, platform engineering, infrastructure engineering, observability engineering, or related disciplines.
  • Strong experience building and operating observability platforms in cloud-native, distributed environments.
  • Deep hands-on experience with several of the following technologies: Prometheus, Thanos, VictoriaMetrics, Grafana, Loki, Tempo, OpenTelemetry, ClickHouse, Elastic, or similar platforms.
  • Strong software engineering skills with proficiency in Go, Python, or equivalent languages.
  • Experience operating and troubleshooting Kubernetes-based platforms at scale.
  • Strong understanding of monitoring, logging, tracing, telemetry pipelines, and modern observability practices.
  • Experience designing systems with scalability, reliability, performance, and operational simplicity in mind.
  • Proficiency with Infrastructure-as-Code tools such as Terraform, Ansible, or equivalent.
  • Ability to lead technical initiatives and influence engineering decisions across multiple teams.
  • Excellent communication skills with the ability to explain technical tradeoffs and align stakeholders around pragmatic solutions.
Preferred
  • Experience operating observability systems in GPU, AI/ML, HPC, or large-scale compute environments.
  • Familiarity with Slurm, Kubernetes GPU scheduling, or AI infrastructure platforms.
  • Experience with high-volume telemetry pipelines and streaming technologies such as Kafka, Vector, or Fluent Bit.
  • Knowledge of observability challenges related to model training, inference workloads, GPU utilization, and distributed AI systems.
  • Experience mentoring engineers and helping grow technical capability across teams.