1

Ml Inference Jobs in Canton, GA (NOW HIRING)

Machine Learning Engineer

Atlanta, GA · On-site

$120 - $165/hr

You will work closely with AI/ML researchers, data engineers, and product teams to design ... Hands‑on experience with LLMs/SLMs (fine‑tuning, prompt design, inference optimization)

... ML model serving patterns (batch vs. real-time inference) for database-adjacent workloads Drive prompt engineering best practices for database-related AI applications

Sr. Director Data & AI Platform Architect

Atlanta, GA · On-site +1

$64.75 - $86.50/hr

Proven ability toarchitectend-to-end ML systems, includingdata pipelines, feature engineering ... inference optimization. US PERSON REQUIREMENTS Due to compliance with U.S. export control laws and ...

Sr. Director Data & AI Platform Architect

Atlanta, GA · On-site +1

$64.75 - $86.50/hr

Proven ability toarchitectend-to-end ML systems, includingdata pipelines, feature engineering ... inference optimization. US PERSON REQUIREMENTS Due to compliance with U.S. export control laws and ...

Explore and evaluate new AI/ML techniques, tools, and methodologies, applying relevant innovations ... and inference efficiency to minimize cost and latency while preserving accuracy. * MLOps ...

Sr Data Engineer

Atlanta, GA

$110K - $132K/yr

... ML workloads * Build and maintain data pipelines for AI product lifecycle, including training data preparation, feature engineering, and inference data flows * Develop and optimize RAG (Retrieval ...

Sr Data Engineer

Atlanta, GA · On-site

$110K - $132K/yr

... ML workloads * Build and maintain data pipelines for AI product lifecycle, including training data preparation, feature engineering, and inference data flows * Develop and optimize RAG (Retrieval ...

Causal inference, experimentation & uplift modelling * NLP, generative AI & multimodal ML systems * Computer vision & video intelligence pipelines * Apply rigorous statistical thinking ...

Sr Data Engineer

Atlanta, GA

$110K - $132K/yr

... ML workloads * Build and maintain data pipelines for AI product lifecycle, including training data preparation, feature engineering, and inference data flows * Develop and optimize RAG (Retrieval ...

Senior AI Engineer

Atlanta, GA · On-site

$100K - $138K/yr

Tune inference performance through KV cache management, paged attention, batching strategies, and ... Build and maintain container images, registries, and CI/CD pipelines for AI/ML services.

The features this framework produces - ML- and LLM-generated alike - power everything from analytics to training to inference to user-facing rendering. What You'll Do * Design and own infrastructure ...

... ML systems. * Proficiency in Python and libraries such as PyTorch, TensorFlow, Scikit-Learn, Hugging Face Transformers. * Hands-on experience with LLMs/SLMs (fine-tuning, prompt design, inference ...

Showing results 41-60

Ml Inference information

See Canton, GA salary details

$35.4K

$115.9K

$185.5K

How much do ml inference jobs pay per year?

As of Aug 15, 2026, the average yearly pay for ml inference in Canton, GA is $115,888.00, according to ZipRecruiter salary data. Most workers in this role earn between $93,000.00 and $128,400.00 per year, depending on experience, location, and employer.

What is ML inference?

ML inference refers to the process of using a trained machine learning model to make predictions or decisions based on new data. After a model has been trained on historical data, inference is the phase where that model is deployed and used in real-world applications, such as recognizing speech, detecting objects in images, or recommending products. The focus in ML inference is on speed, efficiency, and scalability to ensure quick predictions, often in real time. This process is critical for practical applications like mobile apps, web services, and embedded systems. Optimizing inference involves reducing latency, memory usage, and computational requirements.

What is the difference between Ml Inference vs Data Scientist?

AspectML InferenceData Scientist
Required CredentialsKnowledge of machine learning models, programming skillsDegree in data science, statistics, or related fields
Work EnvironmentDeploying models in production, real-time data processingData analysis, model development, research
Industry UsageAI product deployment, software companiesResearch institutions, tech firms, consulting

ML Inference focuses on deploying trained models to make predictions on new data, often in real-time. Data Scientists develop and analyze models, working primarily in research and development. While both roles require understanding of machine learning, ML Inference emphasizes deployment and operationalization, whereas Data Scientists focus on model creation and analysis.

What are some common challenges faced by ML inference engineers when deploying models to production?

ML Inference Engineers often encounter challenges such as optimizing model latency and throughput to meet production requirements, ensuring compatibility with diverse hardware environments, and managing model versioning and updates without disrupting service. Additionally, balancing resource utilization and inference accuracy while monitoring real-time performance metrics is crucial. Collaboration with data scientists, DevOps, and software engineers is typically essential to streamline deployment and maintain robust, scalable inference pipelines.

What are the key skills and qualifications needed to thrive in ML inference?

To thrive in ML Inference, you need a solid background in machine learning principles, programming (Python or C++), and experience with deploying models at scale, often supported by a degree in computer science or a related field. Familiarity with frameworks and tools such as TensorFlow, PyTorch, ONNX, and cloud platforms like AWS SageMaker or Google AI Platform is typically required. Strong problem-solving skills, attention to detail, and effective communication are crucial soft skills for collaborating with multidisciplinary teams and optimizing model performance. These skills ensure efficient, scalable, and reliable deployment of machine learning solutions in real-world applications.

Is ML inference a high paying job?

ML inference roles are generally well-paying, especially for those with skills in machine learning frameworks, programming, and cloud platforms. Salaries vary based on experience, location, and industry, but they tend to be higher than average for tech-related positions.

What are popular job titles related to Ml Inference jobs in Canton, GA?

For Ml Inference jobs in Canton, GA, the most frequently searched job titles are:

What cities near Canton, GA are hiring for Ml Inference jobs?

Cities near Canton, GA with the most Ml Inference job openings:

Senior Software Engineer Applied AI

Advanced Monitored Caregiving Inc.

Atlanta, GA • On-site

$140 - $190/hr

Other

Re-posted yesterday


Job description

Senior Software Engineer: Applied AI (Voice Agents & ML Systems) The pitch

We build and operate production AI voice agents that hold real phone conversations in a regulated healthcare setting, plus the machine learning and LLM pipelines around them. This is one seat that spans four disciplines that rarely come together: real‑time systems, LLM engineering, traditional machine learning, and serious cloud infrastructure, all in production, all with real consequences. If you are the kind of engineer who gets restless doing one thing, this role is the opposite problem.

What you’ll work across
  • Streaming, low‑latency speech‑to‑speech systems built on modern LLMs
  • Telephony and real‑time media (call control, live audio streaming)
  • Audio handling and the quirks of real human conversation (interruptions, timing, noise)
  • Concurrency on a latency‑sensitive path, where p99 matters and a stall is something a caller hears
  • Wrapping nondeterministic models in deterministic control so they behave reliably in production
  • Multi‑model pipelines, prompt design, and cost/latency budgeting
  • Evaluation harnesses, including LLM‑as‑judge and automated agent‑tests‑agent approaches
  • Agentic tooling that gives AI systems safe, structured access to infrastructure
Traditional (non‑LLM) machine learning
  • End‑to‑end ML pipelines: feature engineering, model training, and scheduled inference
  • Imbalanced, messy real‑world data; calibration and explainability for non‑technical consumers
  • Turning research notebooks into reproducible, auditable production pipelines
Cloud and infrastructure
  • Infrastructure as code across multiple environments (we run on AWS)
  • Managed compute, data, streaming, and orchestration services
  • Security engineering in a regulated setting: encryption, least‑privilege access, strict data‑handling discipline
  • Observability and telemetry‑driven debugging, tracing a production issue from a metric anomaly to root cause
Plus

Occasional full‑stack work on internal tools, and an engineering workflow that leans heavily on AI coding assistants, with human accountability for every change.

What you’ll actually do
  • Ship and debug code on a live, real‑time voice pipeline where latency and correctness are user‑facing
  • Design control systems around LLMs: guardrails, budgets, watchdogs, safe fallbacks
  • Build and operate LLM evaluation and batch‑analysis pipelines
  • Own traditional ML workflows from data to scheduled production inference
  • Trace production issues from a metric anomaly to root cause, including building the evidence when the cause is a vendor
Must‑haves
  • 7+ years building and operating production backend systems, with strong general‑purpose programming skills (we work primarily in Python)
  • Experience running distributed systems in the cloud; comfortable debugging from telemetry to root cause
  • Hands‑on production experience with LLMs or generative AI (any provider or framework), plus the judgment to know when not to use a model
  • Working fluency across the traditional machine learning lifecycle (you productionize; you do not need to publish)
  • Disciplined in a regulated environment: small, reviewable changes and careful handling of sensitive data
Nice‑to‑haves
  • Real‑time media or telephony experience
  • Front‑end / full‑stack ability
  • ML pipeline experience, vector search, or embeddings
  • Fluency with AI coding assistants (our workflows assume them, with human accountability for every change)
How we work

Smallest correct change wins. Every behavior change is validated against the live system. Evidence over opinion in debugging. Code review is rigorous. Safety and privacy gate everything.

This role is open only to US citizens and lawful permanent residents (Green Card holders). We cannot consider candidates who require visa sponsorship now or in the future, and we are unable to make exceptions of any kind.

How to apply
  • Your LinkedIn profile URL
  • A phone number where we can reach you

A resume is welcome but optional; the two items above are required.

#J-18808-Ljbffr