1

Ml Infrastructure Jobs in Georgia (NOW HIRING)

GA · On-site

$16.25 - $19.25/hr

Build the evaluation and experimentation infrastructure that lets the Leasing team ship ML changes with confidence -- defining what "better" looks like for leasing-specific tasks and owning the ...

GA · On-site

$16.25 - $19.25/hr

Build the evaluation and experimentation infrastructure that lets the Leasing team ship ML changes with confidence -- defining what "better" looks like for leasing-specific tasks and owning the ...

DevOps/MLOps Engineer

Cumming, GA · On-site

$47 - $64.50/hr

Strong scripting ability in languages like Bash, Python, or TypeScript * 1+ years of experience supporting AI/ML infrastructure, including GPU-based workloads or model deployment * Familiarity with ...

DevOps Platform Engineer

Duluth, GA · On-site

$120 - $180/hr

AI/ML infrastructure experience -- has deployed LLM-based applications or ML models to production; experience with vector databases, model serving APIs, and LLM cost management * CI/CD pipeline ...

ML Software Engineering Lead

Atlanta, GA · On-site

$98K - $129K/yr

Partner closely with research-focused data science teams, business stakeholders, infrastructure ... Strong understanding of the data science/ML research process. * Strong understanding of software ...

AI / ML Engineer

Sandy Springs, GA · On-site

$60 - $70/hr

Business Consultant - AI/ML Engineer Location: Atlanta/Chicago Employment Type: Contract/Permanent ... We are one of the leading providers of digital and AI infrastructure in the world. NTT DATA is a ...

AI/ML Engineer

Atlanta, GA · On-site

$110K - $132K/yr

Infrastructure: Leverage AWS AI/ML services (Sagemaker, Bedrock,Lambda, Step Functions, ECS/EKS) for scalable AI solutions.Data Engineering with PySpark: Optimize large-scale ETL workflows, data ...

next page

Showing results 1-20

Ml Infrastructure information

What is ML infrastructure?

ML Infrastructure refers to the underlying systems, tools, and processes that enable the development, deployment, and scaling of machine learning models. This includes data storage and management, computing resources, model training and serving environments, monitoring, and automation tools. ML Infrastructure ensures that data scientists and engineers can efficiently build, test, and maintain machine learning applications in a reliable and reproducible manner. It is a crucial foundation for organizations looking to operationalize AI and machine learning solutions at scale.

What are some common challenges faced by professionals working in ML infrastructure roles?

Professionals in ML Infrastructure often encounter challenges related to scaling systems to handle large volumes of data, ensuring reliable deployment pipelines, and maintaining reproducibility across different environments. They must also collaborate closely with data scientists and engineers to streamline workflows and address issues like version control and model monitoring. Staying updated with rapidly evolving tools and best practices is essential, and balancing stability with innovation is a frequent aspect of the role.

What are the key skills and qualifications needed to thrive as an ML infrastructure engineer, and why are they important?

To thrive as an ML Infrastructure Engineer, you need a strong background in software engineering, cloud computing, and machine learning concepts, often supported by a degree in computer science or a related field. Proficiency with containerization tools (like Docker and Kubernetes), cloud platforms (such as AWS, GCP, or Azure), and CI/CD systems is critical. Excellent problem-solving, collaboration, and communication skills help you efficiently work with data scientists and DevOps teams. These skills and qualities are vital for building scalable, reliable ML systems that support rapid experimentation and deployment in production environments.

What is the difference between Ml Infrastructure vs Data Engineer?

AspectML InfrastructureData Engineer
Required CredentialsBachelor's in CS, Data Science, or related; knowledge of cloud platformsBachelor's in CS, Software Engineering, or related; experience with databases and ETL tools
Work EnvironmentFocus on deploying and maintaining ML systems, cloud environments, and infrastructure toolsDesigning, building, and managing data pipelines and storage solutions
Industry UsageUsed in AI/ML teams to support model deployment and scalabilityUsed across data-driven organizations for data management and analytics

ML Infrastructure specialists focus on deploying, scaling, and maintaining machine learning systems and infrastructure, while Data Engineers primarily build and manage data pipelines and storage solutions. Both roles require technical skills and often collaborate, but their core responsibilities differ in focus and tools used.

What are popular job titles related to Ml Infrastructure jobs in Georgia?

For Ml Infrastructure jobs in Georgia, the most frequently searched job titles are:

What job categories do people searching Ml Infrastructure jobs in Georgia look for?

The top searched job categories for Ml Infrastructure jobs in Georgia are:

What cities in Georgia are hiring for Ml Infrastructure jobs?

Cities in Georgia with the most Ml Infrastructure job openings:

Infographic showing various Ml Infrastructure job openings in Georgia as of August 2026, with employment types broken down into 91% Full Time, 7% Part Time, and 2% Contract. Highlights an 82% Physical, 5% Hybrid, and 13% Remote job distribution.

AI/ML Ops & Infrastructure Engineer - Q126

R2 Technologies Corporation

Alpharetta, GA • On-site

$105K - $137K/yr

Full-time

Re-posted just now


Job description

Overview:
Job Title: AI/ML Ops & Infrastructure Engineer
Company: R2 Technologies
Location: Alpharetta, GA (Hybrid / Remote Options Available)
Employment Type: Full-Time / Contractual
About R2 Technologies: R2 Technologies is a Certified Minority Business Enterprise (MBE) headquartered in Alpharetta, GA. With over two decades of experience across global markets, we have built a reputation as a trusted partner for IT staffing excellence and cutting-edge digital product innovation. We are driven by innovation and operate on a simple philosophy: "We deliver what we promise, and we promise only what we can deliver." Beyond providing top-tier IT talent, R2 builds cutting-edge proprietary solutions like SmartEnt-an Enterprise AI & IoT Intelligence Platform utilizing advanced NLP and AI technologies. By partnering closely with our clients, we deliver technology-driven outcomes that are realistic, measurable, and impactful.
Job Summary: The shift from classical Machine Learning to Generative AI requires a new breed of infrastructure engineering. R2 Technologies is looking for an AI/ML Ops & Infrastructure Engineer to build and manage the operational backbone for our advanced LLM and agentic systems. You will transition beyond basic CI/CD to implement full-lifecycle LLMOps-managing foundation models, fine-tuned adapters, routing logic, and guardrails. Your work will ensure that our AI solutions, including SmartEnt, run with high performance, optimal GPU utilization, and rigorous compliance.
Key Responsibilities: * Design and maintain highly scalable LLMOps pipelines for continuous integration, evaluation, and deployment of machine learning models and AI agents.
  • Deploy and manage containerized AI applications and model inference servers (e.g., vLLM, Ray Serve, NVIDIA Triton) on Kubernetes across multi-cloud environments (AWS, GCP, Azure).
  • Implement comprehensive observability and trace-level logging for multi-step agentic workflows using platforms like LangSmith, W&B Weave, or MLflow.
  • Automate infrastructure provisioning and monitoring using tools like Terraform and agent-driven workflows (e.g., n8n, GitHub Actions).
  • Optimize GPU computing costs, latency, and token usage for high-traffic AI inference endpoints.
  • Enforce security guardrails, toxic output filtering, and robust access policies within the AI deployment infrastructure.
  • Actively utilize AI-assisted coding tools (Copilot, Cursor) to automate infrastructure-as-code (IaC) and streamline Kubernetes management.

Qualifications: *3 years of hands-on experience in MLOps, DevOps, Site Reliability Engineering (SRE), or Cloud Infrastructure.
  • Strong proficiency in containerization and orchestration (Docker, Kubernetes).
  • Experience with ML/LLM operational platforms (MLflow, Weights & Biases, Databricks Mosaic AI, or SageMaker).
  • Familiarity with serving open-source or fine-tuned LLMs and optimizing inference performance.
  • Proven experience or strong familiarity working alongside AI coding assistants to enhance productivity.
  • Scripting/programming skills in Python and bash, along with experience in CI/CD automation.
  • Passion for the evolving landscape of AI infrastructure, cost-optimization (FinOps), and system reliability.

Skills:
Nvidia,Infrastructure