1

Ai Reliability Engineer Jobs in Winder, GA (NOW HIRING)

Staff Reliability Engineer

Atlanta, GA ยท On-site

$150 - $200/hr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

GCP SRE

Alpharetta, GA ยท On-site

$55.75 - $74/hr

This role is for a Site Reliability Engineer (SRE) with a strong emphasis on Google Cloud Platform (GCP) and AI skills. The position requires robust incident management capabilities within the GCP ...

Staff Reliability Engineer

Atlanta, GA ยท On-site

$165K - $218K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

Staff Reliability Engineer

Atlanta, GA ยท On-site

$98K - $124K/yr

Anduril's family of systems is powered by Lattice OS, an AI-powered operating system that turns ... ABOUT THE TEAM The Reliability Engineering team partners across Anduril's engineering ...

Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

We are seeking a Site Reliability Engineer (SRE) to support a large-scale system modernization and ... Apply Today? you agree to receive calls, AI-generated calls, text messages or emails from Kforce ...

Senior Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

Experience with AI-assisted observability or operational automation * Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms REQUIRED KNOWLEDGE, SKILLS, OR ...

Senior Site Reliability Engineer

Atlanta, GA ยท On-site

$54.75 - $72.75/hr

About Your Role As a Senior Site Reliability Engineer, you will work with our Infrastructure and ... Experience using AI-assisted coding tools (e.g., Claude, GitHub Copilot) to accelerate IaC ...

Site Reliability Engineer I

Atlanta, GA ยท On-site

$98K - $148K/yr

PD) is the global leader in AI-first digital operations. By automatically detecting, diagnosing ... As a Site Reliability Engineer I on the Core Infrastructure team in our Atlanta office, you'll help ...

Showing results 21-40

Ai Reliability Engineer information

See Winder, GA salary details

$53.8K

$104.1K

$124.4K

How much do ai reliability engineer jobs pay per year?

As of Sep 8, 2026, the average yearly pay for ai reliability engineer in Winder, GA is $104,075.00, according to ZipRecruiter salary data. Most workers in this role earn between $90,400.00 and $113,800.00 per year, depending on experience, location, and employer.

What is an AI reliability engineer?

AI Reliability Engineers are professionals responsible for ensuring that artificial intelligence systems function reliably, safely, and effectively over time. They work on monitoring AI models in production, identifying and mitigating potential failures, and improving the robustness of AI systems. Their tasks often include testing, validation, performance monitoring, and implementing best practices for maintaining AI infrastructure. By focusing on reliability, they help organizations deploy AI solutions that are dependable and trustworthy in real-world environments.

What are some common challenges AI reliability engineers face when ensuring model robustness in production environments?

Ai Reliability Engineers often encounter challenges such as monitoring AI model performance for drift or unexpected behavior, managing data quality issues, and implementing automated alerting systems for anomalies. In production, it's crucial to ensure that AI models operate consistently and remain reliable under varying conditions and data inputs. Collaborating closely with data scientists, software engineers, and DevOps teams is essential to address these challenges and to continuously improve model reliability and uptime.

What are the key skills and qualifications needed to thrive as an AI reliability engineer, and why are they important?

To thrive as an AI Reliability Engineer, you need a solid background in computer science or engineering, expertise in AI/ML concepts, and experience with software testing and reliability methodologies. Familiarity with tools like TensorFlow, PyTorch, CI/CD pipelines, and reliability testing frameworks, along with certifications in cloud platforms (e.g., AWS Certified Machine Learning), is highly valuable. Analytical thinking, problem-solving abilities, and strong collaboration skills set top performers apart in this role. These skills ensure robust, dependable AI systems that meet performance standards and maintain trust in critical applications.

What is the difference between Ai Reliability Engineer vs Data Scientist?

AspectAi Reliability EngineerData Scientist
Required CredentialsBachelor's or master's in CS, engineering, or related; certifications in AI/MLBachelor's or master's in CS, statistics, or related; certifications in data analysis or ML
Work EnvironmentTech companies, AI-focused teams, engineering departmentsResearch labs, tech firms, analytics teams
Employer & Industry UsageAI product development, machine learning systems, reliability testingData analysis, predictive modeling, business insights

While both roles involve AI and ML, Ai Reliability Engineers focus on ensuring AI system robustness and uptime, whereas Data Scientists analyze data to generate insights and models. The roles often collaborate but serve different primary functions within AI projects.

What are popular job titles related to Ai Reliability Engineer jobs in Winder, GA?

For Ai Reliability Engineer jobs in Winder, GA, the most frequently searched job titles are:

What cities near Winder, GA are hiring for Ai Reliability Engineer jobs?

Cities near Winder, GA with the most Ai Reliability Engineer job openings:

Senior Databricks AI Platform SRE

Central Business Solutions

Alpharetta, GA โ€ข On-site

$55.75 - $74/hr

Full-time

Re-posted 3 days ago


Job description

Senior Databricks AI Platform SRE
Description: We are looking for a Senior Databricks AI Platform SRE to join our Platform SRE team. This role will be critical in designing, building, and optimizing a scalable, secure, and developer-friendly Databricks platform to enable Machine Learning (ML) and Artificial Intelligence (AI) workloads at enterprise scale
You will partner with ML engineer, data scientists, platform teams, and cloud architects to automate infrastructure, enforce best practices, and streamline the end-to-end ML lifecycle using modern cloud-native technologies.
Total Experience - 5+ Years. Bachelor's or master's degree in computer science, Engineering or a related field.
Responsibilities:
  • Design and implement secure, scalable, and automated Databricks environments to support AI/ML workloads.
  • Develop infrastructure-as-code (IaC) solutions using Terraform for provisioning Databricks, cloud resources, and network configurations.
  • Build automation and self-service capabilities using Python, Java and APIs for platform onboarding, workspace provisioning, orchestration and monitoring.
  • Collaborate with data science and ML teams to define compute requirements, governance policies, and efficient workflows across dev/qa/prod environments.
  • Integrate Databricks offering with cloud-native services on Azure/AWS
  • Champion CI/CD and GitOps for managing ML infrastructure and configurations.
  • Ensure compliance with enterprise security and data governance policies using RBAC, Audit Controls, Encryption, Network Isolation, and policies.
  • Monitor platform performance, reliability, and usage, and drive improvements to optimize cost and resource utilizations
Required Skills:
  • Proven experience with Terraform for building and managing infrastructure.
  • Strong programming skills in Python and Java
  • Hands-on experience with cloud networking, identity and access management, key vaults, monitoring, and logging in Azure
  • Hands on experience with Databricks (Workspace management, Clusters, Jobs, MLFlow, Delta Lake, Unity Catalog, Mosaic AI)
  • Deep understanding of Azure or AWS infrastructure (e.g. IAM, VNets/VPC, Storage, Networks, Compute, Key management, monitoring)
  • Strong experience in distributed system design, development and deployment using agile/devops practices.
  • Experience with CI/CD pipelines (GitHub Actions, or similar)
  • Experience implementing monitoring and observability using Prometheus, Grafana or Databricks-native solutions.
  • Good communication skills, excellent teamwork experience, ability to mentor and develop more junior developers, including participating in constructive code reviews
Preferred Skills:
  • Experience in multi-cloud environments (AWS/GCP) is a bonus
  • Experience in working in highly regulated environments (finance, healthcare, etc.) is desirable
  • Experience with Databricks REST APIs and SDKs
  • Knowledge of MLFlow, Mosaic AC, & MLOps tooling
  • Working with teams using Scrum, Kanban or other agile practices
  • Proficiency with standard Linux command line and debugging tools
  • Azure or AWS Certifications

Central Business Solutions, Inc(A Certified Minority Owned Organization)
Checkout our excellent assessment tool: http://www.skillexam.com/
Checkout our job board : http://www.job-360.net/
Central Business Solutions, Inc
37600 Central Court Suite 214 Newark CA, 94560
Phone: (833)247-8800 Fax: (510)-740-3677
Web: http://www.cbsinfosys.com