1

Ai Reliability Engineer Jobs in Tennessee (NOW HIRING)

Site Reliability Engineer 2

Nashville, TN

$55 - $73.25/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

And with AI embedded across our products and services, we help customers turn that promise into a ... Three or more years of experience in site reliability engineering, systems administration ...

Site Reliability Engineer 2

Nashville, TN · On-site

$55 - $73.25/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

And with AI embedded across our products and services, we help customers turn that promise into a ... Three or more years of experience in site reliability engineering, systems administration ...

Site Reliability Engineer 2

Nashville, TN · On-site

$55 - $73.25/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

And with AI embedded across our products and services, we help customers turn that promise into a ... Three or more years of experience in site reliability engineering, systems administration ...

Site Reliability Engineer 2

Nashville, TN · On-site

$55 - $73.25/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

And with AI embedded across our products and services, we help customers turn that promise into a ... Three or more years of experience in site reliability engineering, systems administration ...

Senior Site Reliability Engineer

Nashville, TN

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

And with AI embedded across our products and services, we help customers turn that promise into a ... Three or more years of experience in site reliability engineering, systems administration ...

Site Reliability Engineer 2

Nashville, TN

$55 - $73.25/hr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

And with AI embedded across our products and services, we help customers turn that promise into a ... Three or more years of experience in site reliability engineering, systems administration ...

Senior Site Reliability Engineer

Nashville, TN · On-site

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

And with AI embedded across our products and services, we help customers turn that promise into a ... Three or more years of experience in site reliability engineering, systems administration ...

Showing results 41-60

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Tennessee?

For Ai Reliability Engineer jobs in Tennessee, the most frequently searched job titles are:

What job categories do people searching Ai Reliability Engineer jobs in Tennessee look for?

The top searched job categories for Ai Reliability Engineer jobs in Tennessee are:

What cities in Tennessee are hiring for Ai Reliability Engineer jobs?

Cities in Tennessee with the most Ai Reliability Engineer job openings:

Senior Manager, Reliability Engineering- (Nashville, TN - onsite)

Oracle Corporation

Nashville, TN • On-site

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

Re-posted 2 days ago


Oracle rating

8.7

Company rating: 8.7 out of 10

Based on 151 frontline employees who took The Breakroom Quiz

55th of 244 rated software companies


Job description

>>This position will be full-time on-site at Oracle's offices located in Nashville, TN.<<  Relocation assistance may be available in accordance with Oracle's relocation policies.

As Senior Manager - Reliability Engineering, you will lead the teams, methods, and programs responsible for improving the availability, maintainability, and lifecycle performance of mission-critical facilities infrastructure across OCI's data center portfolio. This role sets the direction for reliability engineering practices across electrical, mechanical, and controls domains, with a strong focus on analytics, predictive maintenance, risk reduction, and standardized reliability methods.

You will lead engineers and analysts who partner closely with site operations, design, construction, commissioning, and automation teams to identify reliability risks, improve maintenance strategies, strengthen incident learning, and ensure corrective actions are implemented and sustained. This role translates operational data and engineering analysis into portfolio-level standards, priorities, and decisions that protect uptime and support long-term capacity growth.

Only Oracle brings together the data, infrastructure, applications, and expertise to power everything from industry innovations to life-saving care. And with AI embedded across our products and services, we help customers turn that promise into a better future for all. Discover your potential at a company leading the way in AI and cloud solutions that impact billions of lives.

True innovation starts when everyone is empowered to contribute. That's why we're committed to growing a workforce that promotes opportunities for all with competitive benefits that support our people with flexible medical, life insurance, and retirement options. We also encourage employees to give back to their communities through our volunteer programs.

We're committed to including people with disabilities at all stages of the employment process. If you require accessibility assistance or accommodation for a disability at any point, let us know by emailing accommodation-request_mb@oracle.com or by calling 1-888-404-2494 in the United States.

Oracle is an Equal Employment Opportunity Employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, sexual orientation, gender identity, disability and protected veterans' status, or any other characteristic protected by law. Oracle will consider for employment qualified applicants with arrest and conviction records pursuant to applicable law.

Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $126,200 to $264,100 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.

Career Level - M3


Key Responsibilities

  • Lead reliability engineering and analytics teams supporting multiple sites, regions, or infrastructure programs across mission-critical facilities environments. 

  • Define and standardize reliability engineering methods, including FMEA, RCA, RCM, critical assessment, defect elimination, and continuous improvement practices. 

  • Establish reliability programs for electrical, mechanical, and controls systems that improve asset health, reduce repeat failures, and strengthen operational resilience. 

  • Oversee deployment and effective use of monitoring, alarming, analytics, and automation tools that support predictive maintenance, asset performance visibility, and reliability decision-making. 

  • Define, track, and report portfolio reliability of KPIs, trends, risks, and corrective action status to executive leadership and global operations stakeholders. 

  • Ensure corrective actions from incidents, failure analyses, audits, and trend reviews are implemented, verified for effectiveness, and sustained over time. 

  • Partner with site operations teams to improve maintenance programs, troubleshooting rigor, failure response, operational readiness, and procedure quality. 

  • Guide reliability input into commissioning, acceptance testing, operational handover, major maintenance, retrofits, and equipment lifecycle planning. 

  • Drive structured analysis of asset utilization, failure history, equipment health, remaining useful life, spare parts strategy, and end-of-life risk. 

  • Partner with design, construction, and procurement teams to improve reliability, maintainability, standardization, and total cost of ownership in new and existing infrastructure. 

  • Develop engineers and analysts through coaching, technical review, prioritization discipline, and data-driven problem solving. 

  • Build strong operating rhythms across the team, ensuring priorities, escalation paths, portfolio visibility, and follow-through are clear and consistent. 

Ideal Candidate Profile

  • 10+ years of experience in reliability engineering, maintenance engineering, critical facilities, manufacturing, utilities, industrial operations, or other uptime-critical environments; data center experience preferred, but adjacent industry experience is highly relevant. 

  • 2-4+ years of experience leading engineers, analysts, or technical programs across multiple sites, business units, or complex infrastructure environments. 

  • Strong background in reliability frameworks, structured root cause analysis, maintenance strategy, asset lifecycle planning, and operational risk reduction. 

  • Experience working cross-functionally with operations, design, construction, commissioning, and vendor partners. 

  • Bachelor's degree in Electrical Engineering, Mechanical Engineering, Industrial Engineering, Systems Engineering, OR relevant industry experience

Skills and Competencies

  • Strong technical leadership and ability to influence engineering, operations, and business stakeholders. 

  • Strong analytical judgment and ability to translate reliability data into practical decisions, priorities, and executive-level recommendations. 

  • Deep familiarity with RCA rigor, FMEA/FMECA, reliability metrics, maintenance optimization, and continuous improvement methods. 

  • Ability to connect site-level issues to portfolio-wide patterns, risks, and opportunities. 

  • Clear communicator who can explain technical tradeoffs, operational risk, and recommended actions to mixed audiences. 

  • Strong organizational discipline, follow-through, and ability to manage multiple priorities in a fast-moving environment. 

Preferred Skills

  • Experience in hyperscale data centers or other mission-critical environments supporting electrical distribution, mechanical cooling systems, controls/automation platforms, and integrated facility operations. 

  • Experience with predictive maintenance programs, condition-based monitoring, failure trend analysis, and asset health modeling. 

  • Familiarity with reliability, availability, and maintainability analysis; maintenance strategy development; spare parts optimization; and lifecycle cost evaluation. 

  • Experience supporting or governing commissioning, operational acceptance, maintenance of program design, and system readiness for new or modified installations. 

  • Working knowledge of monitoring and analytics platforms, CMMS/EAM systems, DCIM, EPMS/BMS data, and automation tools used to improve facility reliability. 

  • Experience building KPI dashboards and reporting frameworks for senior operations or executive leadership. 

  • Knowledge of reliability practices across electrical, mechanical, and controls systems, even if prior domain depth is strongest in one area. 

Preferred Credentials / Certifications

  • Certified Maintenance & Reliability Professional (CMRP) preferred. 

  • Certified Reliability Engineer (CRE) preferred. 

  • ASQ, SMRP, or equivalent reliability and quality certifications are a plus. 

  • Data center or critical environment credentials such as Uptime Institute training/certifications are a plus where relevant to the role. 

  • OEM, controls, analytics, or condition-monitoring training relevant to critical infrastructure reliability programs is a plus. 

  • Advanced training in FMEA, RCA, Lean, Six Sigma, or structured problem-solving methodologies is a plus. 

Eligibility and Location Requirements
This position requires U.S. citizenship and is located onsite in Nashville, TN. Relocation assistance may be available in accordance with Oracle's relocation policies.

Why Oracle Cloud Infrastructure?

Global impact at scale: Contribute directly to how mission-critical OCI data centers operate across regions and continents, influencing infrastructure reliability, security, sustainability, and long-term capacity growth.

Technically rigorous environment: Work alongside experienced engineers, automation specialists, and compliance teams in a rapidly scaling hyperscale cloud infrastructure, where disciplined execution and technical depth matter.

Culture built on operational excellence: Join an organization that values safety, process rigor, clear accountability, and continuous improvement as foundational to protecting uptime and customer trust.

Long-term career development: Benefit from internal mobility, role-based technical training, and development opportunities designed for professionals building long-term careers in cloud infrastructure and facilities operations.

#LI-SB36


What Oracle employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Oracle logo

About Oracle

Sourced by ZipRecruiter

An Oracle career can span industries, roles, Countries and cultures, giving you the opportunity to flourish in new roles and innovate, while blending work life in. Oracle has thrived through 40+ years of change by innovating and operating with integrity while delivering for the top companies in almost every industry. In order to nurture the talent that makes this happen, we are committed to an inclusive culture that celebrates and values diverse insights and perspectives, a workforce that inspires thought leadership and innovation. Oracle offers a highly competitive suite of Employee Benefits designed on the principles of parity, consistency, and affordability. The overall package includes certain core elements such as Medical, Life Insurance, access to Retirement Planning, and much more. We also encourage our employees to engage in the culture of giving back to the communities where we live and do business. At Oracle, we believe that innovation starts with diversity and inclusion and to create the future we need talent from various backgrounds, perspectives, and abilities. We ensure that individuals with disabilities are provided reasonable accommodation to successfully participate in the job application, interview process, and in potential roles. to perform crucial job functions. That's why we're committed to creating a workforce where all individuals can do their best work. It's when everyone's voice is heard and valued that we're inspired to go beyond what's been done before.

Industry

It services

Company size

10,000+ Employees

Headquarters location

Redwood City, CA, US

Year founded

1977

Social media