1

Ai Reliability Engineer Jobs in Nevada (NOW HIRING)

Senior Machine Learning Engineer

Las Vegas, NV · On-site

$117K - $154K/yr

Because breakthrough AI should move at the speed of ideas, not infrastructure. About the Role We're ... Troubleshoot performance, reliability, and scalability issues across the ML stack * Partner with ML ...

Senior Machine Learning Engineer

Las Vegas, NV · On-site +1

$117K - $154K/yr

Because breakthrough AI should move at the speed of ideas, not infrastructure. About the Role We're ... Troubleshoot performance, reliability, and scalability issues across the ML stack * Partner with ML ...

Sr Data Engineer

Las Vegas, NV · On-site

$109K - $131K/yr

... reliability • Design and evolve data architecture supporting analytics, BI, and future AI/ML ... Required : • Bachelor's degree required in Computer Science, Data Science, Engineering ...

NPI Hardware Engineer Lead

Las Vegas, NV · On-site

$97K - $128K/yr

... AI, cloud, and connected infrastructure. As a US-based manufacturing partner, the company rapidly ... Experience with tolerance analysis, DFMA, and mechanical reliability testing. * Knowledge of test ...

NPI Hardware Engineer Lead

Las Vegas, NV · On-site

$118K - $155K/yr

... AI, cloud, and connected infrastructure. As a US-based manufacturing partner, the company rapidly ... Experience with tolerance analysis, DFMA, and mechanical reliability testing. * Knowledge of test ...

NPI Hardware Engineer

Las Vegas, NV · On-site

$118K - $155K/yr

... AI, cloud, and connected infrastructure. As a US-based manufacturing partner, the company rapidly ... Experience with tolerance analysis, DFMA, and mechanical reliability testing. * Knowledge of test ...

Data Engineer II

Las Vegas, NV

$110K - $132K/yr

... AI initiatives. You will work with modern data technologies including Databricks, Apache Spark ... Monitor, troubleshoot, and optimize production data pipelines for reliability, scalability, and ...

Data Engineer II

Las Vegas, NV

$110K - $132K/yr

... AI initiatives. You will work with modern data technologies including Databricks, Apache Spark ... Monitor, troubleshoot, and optimize production data pipelines for reliability, scalability, and ...

Data Engineer II

Reno, NV

$114K - $137K/yr

... AI initiatives. You will work with modern data technologies including Databricks, Apache Spark ... Monitor, troubleshoot, and optimize production data pipelines for reliability, scalability, and ...

Data Engineer II

Reno, NV

$114K - $137K/yr

... AI initiatives. You will work with modern data technologies including Databricks, Apache Spark ... Monitor, troubleshoot, and optimize production data pipelines for reliability, scalability, and ...

NPI Hardware Engineer

Las Vegas, NV · On-site

$118K - $155K/yr

... AI, cloud, and connected infrastructure. As a US-based manufacturing partner, the company rapidly ... Experience with tolerance analysis, DFMA, and mechanical reliability testing. * Knowledge of test ...

TensorWave is dedicated to delivering seamless, secure, reliable, and resilient AI compute at scale. The Staff Database Engineer will be responsible for database architecture, reliability ...

They understand why systems are designed a certain way, how technical decisions impact reliability ... AI-assisted automation tools What Success Looks Like This role is ideal for someone who enjoys ...

They understand why systems are designed a certain way, how technical decisions impact reliability ... AI-assisted automation tools What Success Looks Like This role is ideal for someone who enjoys ...

Senior Software Engineer

Mccarran, NV · On-site

$125K - $165K/yr

Debug and resolve production issues; monitor and improve system reliability * Collaborate remotely ... Enthusiasm for AI-assisted development * Scrappy mindset: you've built under tight timelines and ...

Showing results 41-60

Ai Reliability Engineer information

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.
What are popular job titles related to Ai Reliability Engineer jobs in Nevada? For Ai Reliability Engineer jobs in Nevada, the most frequently searched job titles are:
What job categories do people searching Ai Reliability Engineer jobs in Nevada look for? The top searched job categories for Ai Reliability Engineer jobs in Nevada are:
What cities in Nevada are hiring for Ai Reliability Engineer jobs? Cities in Nevada with the most Ai Reliability Engineer job openings:
Infographic showing various Ai Reliability Engineer job openings in Nevada as of August 2026, with employment types broken down into 75% Full Time, 21% Part Time, and 4% Contract. Highlights an 72% Physical, 3% Hybrid, and 25% Remote job distribution.

Senior Machine Learning Engineer

TensorWave

Las Vegas, NV • On-site

$117K - $154K/yr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Re-posted 22 days ago


Job description

About TensorWave

Our mission is simple: deliver seamless, secure, reliable, and resilient AI compute at scale. We've built a versatile cloud platform that eliminates infrastructure barriers, empowering builders to focus on innovation instead of fighting their stack. Because breakthrough AI should move at the speed of ideas, not infrastructure.

 

About the Role

We’re looking for a Senior Machine Learning Engineer to join our team during an exciting phase of growth. In this role, you’ll be responsible for building and operating the core systems that power large-scale ML training and inference across TensorWave’s GPU platform, working closely with cross-functional partners to support business objectives while upholding our standards for excellence, collaboration, and impact.

 

What You’ll Do

  • Design, operate, and improve ML infrastructure systems supporting distributed training and inference workloads

  • Build reliable, repeatable workload execution and orchestration patterns across shared GPU environments

  • Troubleshoot performance, reliability, and scalability issues across the ML stack

  • Partner with ML, systems, and platform teams to improve developer experience and operational efficiency

 

Who You Are

Required Qualifications

  • Bachelor of Science in Computer Science, Computer Engineering, or a related technical field, or equivalent practical experience

  • Expertise supporting production ML systems using SLURM and Kubernetes

  • Strong understanding of GPU-accelerated workloads and distributed systems concepts

  • Solid Linux fundamentals and experience debugging infrastructure-level issues

  • Ability to build automation and tooling - Python, Go, etc.

Preferred Qualifications

  • Experience working across schedulers, orchestration platforms, or cluster managers

  • Familiarity with large-scale GPU environments or HPC-style systems

  • Experience improving infrastructure reliability, utilization, or performance at scale

 

What We Offer

  • Stock Options

  • 100% paid Medical, Dental, and Vision insurance for Employees

  • Company Health Savings Account Contributions

  • 100% paid Short Term and Long Term Disability Insurance for Employees

  • Life and Voluntary Supplemental Insurance Options

  • Other Insurance Options, such as Pet & Legal Insurance

  • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support

  • Flexible Spending Account

  • 401(k)

  • Employee Assistance Program

  • Flexible PTO

  • Paid Holidays

  • Parental Leave

  • Other In-Office Perks

 

Equal Employment Opportunity

TensorWave is an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate on the basis of any protected status under applicable law.

 

Reasonable Accommodations

TensorWave provides reasonable accommodations in accordance with applicable laws. If you require accommodation during the hiring process, please contact accomodations@tensorwave.com.

 

Employment Eligibility

All offers of employment are contingent upon verification of identity and authorization to work in the United States, as required by law.

 

Background Checks

Where permitted by law, employment may be contingent upon the successful completion of a job-related background check.

 

Data Privacy Notice

By submitting an application, you acknowledge that TensorWave may collect, use, and retain your personal information for recruiting and employment-related purposes in accordance with applicable data privacy laws.