1

Ai Reliability Engineer Jobs in Minnesota (NOW HIRING)

Sr Engineer - US

Brooklyn Park, MN · On-site

$98K - $176K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... long-term reliability. * Engineer systems that balance high throughput, low latency, resiliency ... Integrate AI-driven capabilities, including risk-aware access identification and contextual access ...

Sr Engineer - US

Brooklyn Park, MN · On-site

$98K - $176K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... long-term reliability. * Engineer systems that balance high throughput, low latency, resiliency ... Integrate AI-driven capabilities, including risk-aware access identification and contextual access ...

New

Senior AI Engineer - Circlecard

Brooklyn Park, MN · On-site

$98K - $176K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... reliability. What You'll Do AI Engineering * Leverage AI-assisted software development tools ... throughout the software development lifecycle, including solution design, implementation, testing ...

New

Senior AI Engineer - Circlecard

Minneapolis, MN · On-site

$98K - $176K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... reliability. What You'll Do AI Engineering * Leverage AI-assisted software development tools ... throughout the software development lifecycle, including solution design, implementation, testing ...

Sr Engineer - US

Minneapolis, MN · On-site +1

$98K - $176K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... long-term reliability. * Engineer systems that balance high throughput, low latency, resiliency ... Integrate AI-driven capabilities, including risk-aware access identification and contextual access ...

AI Engineer

Minneapolis, MN · On-site

$85K - $128K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

AI Platform Engineer SUMMARY The AI Platform Engineer designs, builds, and maintains the ... Performance Monitoring & Reliability * Establish observability for AI workloads (metrics, logs ...

AI Engineer

Minneapolis, MN · On-site

$85K - $128K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

AI Platform Engineer SUMMARY The AI Platform Engineer designs, builds, and maintains the ... Performance Monitoring & Reliability * Establish observability for AI workloads (metrics, logs ...

AI Engineer

Minneapolis, MN · On-site

$85K - $128K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

AI Platform Engineer SUMMARY The AI Platform Engineer designs, builds, and maintains the ... Performance Monitoring & Reliability * Establish observability for AI workloads (metrics, logs ...

Senior Inference Engineer - AI

Eagan, MN · Hybrid

$106K - $146K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Implement observability and health monitoring for inference pipelines, ensuring reliability of enterprise AI services * Collaborates closely with AI engineers to invent new quantization techniques ...

Senior Inference Engineer - AI

Eagan, MN · On-site

$106K - $146K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Implement observability and health monitoring for inference pipelines, ensuring reliability of enterprise AI services * Collaborates closely with AI engineers to invent new quantization techniques ...

New

Senior Inference Engineer - AI

Eagan, MN · On-site

$106K - $146K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Implement observability and health monitoring for inference pipelines, ensuring reliability of enterprise AI services * Collaborates closely with AI engineers to invent new quantization techniques ...

Showing results 41-60

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Minnesota?

For Ai Reliability Engineer jobs in Minnesota, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Minnesota look for?

The top searched job categories for Ai Reliability Engineer jobs in Minnesota are:

What cities in Minnesota are hiring for Ai Reliability Engineer jobs?

Cities in Minnesota with the most Ai Reliability Engineer job openings:

Network Engineer, AI Infrastructure Repair

Meta

Rosemount, MN

$193K/yr

Full-time

Re-posted 12 days ago


Meta rating

7.8

Company rating: 7.8 out of 10

Based on 45 frontline employees who took The Breakroom Quiz

136th of 244 rated software companies


Job description

Meta is building the next generation of AI infrastructure to power large-scale machine learning workloads, and the reliability of that infrastructure depends on reliable, high-performance network engineering. In this role, you will lead the strategy and execution for AI network repair and remediation programs, ensuring that the high-performance fabrics underpinning Meta's AI training and inference clusters remain operational, resilient, and optimized. You will drive cross-functional initiatives spanning network deployment, fault diagnosis, and repair automation across Meta's AI data center environments, shaping the systems and processes that keep AI infrastructure at scale.
Network Engineer, AI Infrastructure Repair Responsibilities:
  • Define and drive the long-term strategy for AI network repair and remediation programs across large-scale data center environments supporting machine learning workloads
  • Lead root cause analysis and resolution of complex network faults affecting high-performance AI training and inference fabrics, including RDMA, high-speed Ethernet, and optical interconnect layers
  • Develop and champion novel approaches to network fault detection, automated remediation, and repair workflow optimization for AI cluster infrastructure
  • Partner with hardware, software, and data center operations teams to align network repair programs with AI infrastructure deployment roadmaps and capacity plans
  • Establish and refine operational frameworks, runbooks, and tooling for network repair at scale, reducing mean time to repair across AI fabric environments
  • Identify systemic reliability risks in AI network infrastructure and drive cross-functional initiatives to address them before they impact production workloads
  • Influence the design of next-generation AI network architectures by contributing repair and reliability insights to hardware and topology decisions
  • Leverage AI-driven analytics and automation tools to redesign repair workflows, accelerating fault identification and resolution across distributed network environments
  • Build and maintain strategic relationships with internal engineering, operations, and vendor partners to ensure repair programs scale with AI infrastructure growth
  • Communicate program status, risk, and strategic recommendations to engineering leaders and cross-functional stakeholders through structured reporting and executive briefings

Minimum Qualifications:
  • Experience influencing technical direction and organizational strategy through data-driven analysis, written proposals, and stakeholder alignment across engineering and operations teams
  • Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
  • Experience leading cross-functional programs that span network operations, hardware deployment, and infrastructure reliability at data center scale
  • Experience developing and driving strategy for network fault management, repair automation, or remediation programs in production environments
  • Experience designing, deploying, or operating high-speed network fabrics used in AI or machine learning infrastructure, including technologies such as RDMA over Converged Ethernet, InfiniBand, or high-density optical interconnects
  • 12+ years of experience in network engineering, with a focus on large-scale data center or high-performance computing network environments

Preferred Qualifications:
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Experience with network telemetry platforms, observability tooling, or AI-assisted anomaly detection applied to large-scale fabric environments
  • Experience building or scaling repair operations programs, including workforce planning, tooling development, and process standardization across multiple data center sites
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
  • Track record of contributing to network hardware or topology design reviews, translating operational repair insights into upstream engineering improvements
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
  • Familiarity with AI accelerator interconnect architectures and the network reliability requirements of distributed training workloads at hyperscale

About Meta:
Meta builds technologies that help people connect, find communities, and grow businesses. When Facebook launched in 2004, it changed the way people connect. Apps like Messenger, Instagram and WhatsApp further empowered billions around the world. Now, Meta is moving beyond 2D screens toward immersive experiences like augmented and virtual reality to help build the next evolution in social technology. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible today—beyond the constraints of screens, the limits of distance, and even the rules of physics.
Meta is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender, gender identity, gender expression, transgender status, sexual stereotypes, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Meta participates in the E-Verify program in certain locations, as required by law. Please note that Meta may leverage artificial intelligence and machine learning technologies in connection with applications for employment.
Meta is committed to providing reasonable accommodations for candidates with disabilities in our recruiting process. If you need any assistance or accommodations due to a disability, please let us know at accommodations-ext@meta.com.
$193,000/year to $271,000/year + bonus + equity + benefits
Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate, monthly rate, or annual salary only, and do not include bonus, equity or sales incentives, if applicable. In addition to base compensation, Meta offers benefits. Learn more about benefits at Meta.

What Meta employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom