1

Ai Reliability Engineer Jobs in Mississippi (NOW HIRING)

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid ... reliability energy systems. * Strong understanding of turbine operations, control systems, and ...

Principal Software Engineer (Python)

Oxford, MS · On-site +1

$127K - $170K/yr

Develop High-Performance AI Services: Utilize Git or similar, Unix command line, Python, FastAPI ... reliability of non-deterministic semantic pipelines. * Apply Advanced Algorithms: Utilize ...

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid ... reliability. Expect to troubleshoot complex issues, streamline processes, and collaborate with ...

Staff Software Engineer

West, MS · On-site

$171 - $252/hr

ACCELERATE WITH AI: Lead by example in using AI coding assistants (GitHub Copilot, Anthropic Claude ... class reliability and security.* WEAR THE CUSTOMER'S SHOES: Partner with Product Managers to ...

System Engineer

Brookhaven, MS · On-site

$71 - $98/hr

What we need Symbotic is seeking a System Engineer to support the day-to-day reliability ... end-to-end, AI-powered robotic and software platform. Symbotic reinvents the warehouse as a ...

What we need Symbotic is seeking a System Engineer to support the day-to-day reliability ... end-to-end, AI-powered robotic and software platform. Symbotic reinvents the warehouse as a ...

What we need Symboticis seeking a System Engineer to support the day-to-day reliability ... end-to-end, AI-powered robotic and software platform. Symbotic reinvents the warehouse as a ...

next page

Showing results 1-20

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Mississippi?

For Ai Reliability Engineer jobs in Mississippi, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Mississippi look for?

The top searched job categories for Ai Reliability Engineer jobs in Mississippi are:

What cities in Mississippi are hiring for Ai Reliability Engineer jobs?

Cities in Mississippi with the most Ai Reliability Engineer job openings:

Site Reliability Engineer - Data Center

SpaceXAI

Southaven, MS

$53.50 - $71.25/hr

Full-time

Posted 4 days ago


SpaceX rating

8.7

Company rating: 8.7 out of 10

Based on 149 frontline employees who took The Breakroom Quiz

17th of 72 rated aerospace companies


Job description

SpaceXAI's mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company's mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates.

ABOUT THE ROLE:

As a Site Reliability Engineer focused on Hardware, you will serve as an expert focused on firmware, hardware specifications, vendor relations, and failure analysis. You will proactively identify and resolve hardware issues, manage RMA processes, and stay ahead of emerging hardware technologies to support SpaceXAI's data center operations. This role demands deep technical expertise in hardware diagnostics, and forward-looking hardware evaluation.

RESPONSIBILITIES:
  • Analyze firmware packages and hardware specifications for upcoming releases for compatibility, performance, and reliability in SpaceXAI's data center environment. Run security scanning and CVE / vulnerability analysis on firmware and related components. Flag safety issues (electrical, thermal, power-protection, fail-safe behavior) before the package hits the floor.
  • Investigate and diagnose hardware failures, including "grey failures" (ambiguous or intermittent issues), proving them as true hardware defects through rigorous testing and data analysis.
  • Manage vendor relationships, including initiating RMA (Return Merchandise Authorization) claims, negotiating beyond standard processes when necessary, and holding vendors accountable for resolutions.
  • Collaborate with Data Center Operations Technicians to troubleshoot, repair, and optimize hardware systems in real-time.
  • Develop and implement monitoring tools, scripts, and processes to detect hardware anomalies early and minimize downtime.
  • Document failure modes, RCAs, AFR / reliability models, RMA outcomes, and hardware evaluations into a team knowledge base.
  • Participate in on-call rotations and incident response for hardware-related issues in the Memphis data center
BASIC QUALIFICATIONS:
  • Bachelor's degree in Systems Engineering, Electrical Engineering, Computer Science, or a related field (or equivalent experience).
  • 2+ years of experience in hardware reliability engineering, preferably in high-performance computing or data center environments.
  • Proven expertise in firmware analysis, hardware specifications review, and release validation.
  • Strong experience with RMA processes, including filing claims, vendor negotiations, and pushing for resolutions outside standard protocols.
  • Demonstrated ability to diagnose and prove complex hardware failures, including grey or intermittent issues, using tools, logic analyzers, or diagnostic software.
  • Familiarity with data center hardware components (e.g., servers, GPUs, networking equipment) and emerging technologies.
  • Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Rust, or similar). Not required to be expert in all of them.
  • Excellent problem-solving skills with a data-driven approach to reliability engineering.
  • Ability to work collaboratively with cross-functional teams, including operations technicians.
PREFERRED SKILLS AND EXPERIENCE:
  • Experience in AI/ML infrastructure or supercomputing environments.
  • Knowledge of vendor ecosystems (e.g., NVIDIA, Dell, HP, Supermicro) and supply chain management.
  • Certifications in hardware engineering or reliability (e.g., CRE, CompTIA Server+).
  • Prior work in a fast-paced startup or tech company like SpaceXAI.

SpaceXAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice.


What SpaceX employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


SpaceX logo

About SpaceX

Sourced by ZipRecruiter

Industry

Aerospace product and parts manufacturing, data services, guided missile and space vehicle manufacturing and satellite telecommunications

Company size

1,001 - 5,000 Employees

Headquarters location

Hawthorne, CA, US