1

Ai Reliability Engineer Jobs in Georgia (NOW HIRING)

Senior Site Reliability Engineer

Atlanta, GA · On-site

$54.75 - $72.75/hr

About Your Role As a Senior Site Reliability Engineer, you will work with our Infrastructure and ... Experience using AI-assisted coding tools (e.g., Claude, GitHub Copilot) to accelerate IaC ...

$46.50 - $61.75/hr

We unify SEO authority and AI visibility, so brands are found, cited, and chosen everywhere search ... About the role We're the SRE Team, the specialists behind the reliability of Semrush's robust ...

Site Reliability Engineer I

Atlanta, GA · On-site

$98K - $148K/yr

PD) is the global leader in AI-first digital operations. By automatically detecting, diagnosing ... As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help ...

Site Reliability Engineer II

Atlanta, GA · On-site

$113K - $171K/yr

PD) is the global leader in AI-first digital operations. By automatically detecting, diagnosing ... As a Site Reliability Engineer II on the Core Infrastructure team in our Atlanta office, you'll ...

Lead Site Reliability Engineer

Alpharetta, GA · On-site

$55.75 - $74/hr

This is a Lead Site Reliability Engineer position at Vice President level, which is part of the job ... Hands-on with AI and implementation of AI tools for operational efficiency * Strong ownership ...

Showing results 41-60

Ai Reliability Engineer information

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What job categories do people searching Ai Reliability Engineer jobs in Georgia look for?

The top searched job categories for Ai Reliability Engineer jobs in Georgia are:

What cities in Georgia are hiring for Ai Reliability Engineer jobs?

Cities in Georgia with the most Ai Reliability Engineer job openings:

Infographic showing various Ai Reliability Engineer job openings in Georgia as of August 2026, with employment types broken down into 45% Full Time, and 55% Contract. Highlights an 95% In-person, and 5% Remote job distribution.

Senior Databricks AI Platform SRE

Central Business Solutions Inc.

Alpharetta, GA • On-site

$54.25 - $72/hr

Full-time

Re-posted 2 days ago


Job description

Job Summary:
Central Business Solutions Inc. is looking for a Senior Databricks AI Platform SRE to join their Platform SRE team. This role will be critical in designing, building, and optimizing a scalable, secure, and developer-friendly Databricks platform to enable Machine Learning (ML) and Artificial Intelligence (AI) workloads at enterprise scale.
Responsibilities:
• Design and implement secure, scalable, and automated Databricks environments to support AI/ML workloads.
• Develop infrastructure-as-code (IaC) solutions using Terraform for provisioning Databricks, cloud resources, and network configurations.
• Build automation and self-service capabilities using Python, Java and APIs for platform onboarding, workspace provisioning, orchestration and monitoring.
• Collaborate with data science and ML teams to define compute requirements, governance policies, and efficient workflows across dev/qa/prod environments.
• Integrate Databricks offering with cloud-native services on Azure/AWS
• Champion CI/CD and GitOps for managing ML infrastructure and configurations.
• Ensure compliance with enterprise security and data governance policies using RBAC, Audit Controls, Encryption, Network Isolation, and policies.
• Monitor platform performance, reliability, and usage, and drive improvements to optimize cost and resource utilizations.
Qualifications:
Required:
• Total Experience – 5+ Years. Bachelor's or master's degree in computer science, Engineering or a related field.
• Proven experience with Terraform for building and managing infrastructure.
• Strong programming skills in Python and Java
• Hands-on experience with cloud networking, identity and access management, key vaults, monitoring, and logging in Azure
• Hands on experience with Databricks (Workspace management, Clusters, Jobs, MLFlow, Delta Lake, Unity Catalog, Mosaic AI)
• Deep understanding of Azure or AWS infrastructure (e.g. IAM, VNets/VPC, Storage, Networks, Compute, Key management, monitoring)
• Strong experience in distributed system design, development and deployment using agile/devops practices.
• Experience with CI/CD pipelines (GitHub Actions, or similar)
• Experience implementing monitoring and observability using Prometheus, Grafana or Databricks-native solutions.
• Good communication skills, excellent teamwork experience, ability to mentor and develop more junior developers, including participating in constructive code reviews
Preferred:
• Experience in multi-cloud environments (AWS/GCP) is a bonus
• Experience in working in highly regulated environments (finance, healthcare, etc.) is desirable
• Experience with Databricks REST APIs and SDKs
• Knowledge of MLFlow, Mosaic AC, & MLOps tooling
• Working with teams using Scrum, Kanban or other agile practices
• Proficiency with standard Linux command line and debugging tools
• Azure or AWS Certifications
Company:
Central Business Solutions Inc.: Your Partner in Technology & Business Transformation CBSInfosys: 25+ years delivering integrated tech/business solutions. Founded in 2000, the company is headquartered in Newark, USA, with a team of 201-500 employees. The company is currently Growth Stage.