1

Machine Learning Data Associate Jobs in Indianapolis, IN

Guides students through data preprocessing, feature selection, building and comparing ... Familiar with machine learning curricula and common challenges such as understanding bias-variance ...

Role Summary The data science (DS) internship at Crowe follows the firmwide calendar, approximately overlapping the academic summer. DS interns will have a designated data scientist mentor and will ...

Deep knowledge of statistical analysis, data wrangling, exploratory data analysis, machine learning, data visualization, SQL, Python or R programming, hypothesis testing, and communication of data ...

Description Launch your machine learning career in Greenfield, IN as a Machine Technical Associate with the ability to grow within the organization! This Temp-to-Hire opportunity pays $22-$30 an hour.

next page

Showing results 1-20

Machine Learning Data Associate information

See Indianapolis, IN salary details

$9

$17

$29

How much do machine learning data associate jobs pay per hour?

As of Sep 12, 2026, the average hourly pay for machine learning data associate in Indianapolis, IN is $17.91, according to ZipRecruiter salary data. Most workers in this role earn between $14.71 and $19.09 per hour, depending on experience, location, and employer.

What is a machine learning data associate?

Machine Learning Data Associates are professionals who support the development of machine learning models by preparing, labeling, and validating data sets. Their work ensures that data used for training algorithms is accurate, consistent, and properly annotated. They may also assist with data cleaning, quality checks, and sometimes basic data analysis tasks. This role is crucial in industries where high-quality labeled data is essential for building effective AI systems.

What are the key skills and qualifications needed to thrive as a machine learning data associate?

To thrive as a Machine Learning Data Associate, you need strong analytical skills, attention to detail, and a basic understanding of data annotation and labeling processes, often supported by a degree in computer science or a related field. Familiarity with data management tools, annotation platforms, and sometimes scripting languages like Python is typically required. Strong communication, collaboration, and problem-solving abilities help you work efficiently with data science teams and ensure high-quality outcomes. These skills and qualities are crucial for producing accurate datasets that directly impact the effectiveness of machine learning models.

How does a machine learning data associate typically collaborate with data scientists and engineers within a project team?

As a Machine Learning Data Associate, you play a vital role in supporting data scientists and engineers by annotating, cleaning, and organizing large datasets to ensure high data quality. You'll frequently communicate with team members to clarify labeling guidelines, provide feedback on data inconsistencies, and report any edge cases encountered during annotation. This collaboration ensures that the datasets used for training machine learning models are accurate and comprehensive, directly impacting the success of the project. Expect regular team meetings and ongoing feedback loops to maintain alignment with evolving project requirements.

What is the difference between Machine Learning Data Associate vs Data Analyst?

AspectMachine Learning Data AssociateData Analyst
Required SkillsData cleaning, labeling, basic programming, understanding of ML workflowsData interpretation, visualization, statistical analysis
Work EnvironmentTech companies, AI startups, research labsBusiness, finance, marketing, healthcare sectors
Common CertificationsData Science certifications, Python, SQLExcel, Tableau, SQL certifications

The main difference is that Machine Learning Data Associates focus on preparing and labeling data specifically for machine learning models, while Data Analysts interpret data to generate insights for business decisions. Both roles require strong data skills and often overlap, but their primary objectives and work environments differ.

How do I become a machine learning data associate?

To become a machine learning data associate, candidates typically need a high school diploma or equivalent, with some roles preferring a bachelor's degree in computer science, data science, or related fields. Relevant skills include data annotation, understanding of machine learning concepts, and proficiency with tools like Excel, SQL, or data labeling platforms. Gaining experience through internships or certifications can improve job prospects in this field.

Is a Machine Learning Data Associate a good job?

A Machine Learning Data Associate role involves preparing and managing data for machine learning models, often requiring skills in data cleaning, annotation, and familiarity with tools like Python or SQL. It can be a good entry-level position for those interested in AI and data science, offering opportunities to develop technical skills and gain industry experience. Compensation and job satisfaction vary depending on the employer and location, but it generally provides a solid foundation for a career in machine learning or data analysis.

What cities near Indianapolis, IN are hiring for Machine Learning Data Associate jobs?

Cities near Indianapolis, IN with the most Machine Learning Data Associate job openings:

Infographic showing various Machine Learning Data Associate job openings in Indianapolis, IN as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 13% Part Time, and 4% Contract. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution, with an average salary of $37,254 per year, or $17.9 per hour.

Machine Learning & Data Operations Engineer

Indianapolis, IN • On-site

Eli Lilly and Company
Pharmaceutical Product Wholesalers • 10K+ employees

$109K - $131K/yr

Full-time

Medical, Dental, Vision, Life, Retirement, PTO

This job post has expired 1 day ago. Applications are no longer accepted.


Eli Lilly and Company rating

8.9

Company rating: 8.9 out of 10

Based on 64 frontline employees who took The Breakroom Quiz


Job description

At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work-but it's work worth doing. If you're driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.


Position Summary

As a Machine Learning & Data Operations Engineer on TuneLab, you will build cutting-edge ML and AI tools alongside a team of engineers and scientists to accelerate and enhance Lilly's drug discovery process. You will take a hands-on role across the full lifecycle of models and the data that feeds them: moving trained models from research into reliable production environments, running inference at scale, and building the data pipelines and readiness checks that keep the data substrate underpinning those models trustworthy. You will stand up the validation, monitoring, and model-card review that keep both models and data production-ready-catching anomalies, schema drift, and performance regressions before they reach researchers. You will collaborate closely with partners across Lilly Research Labs, AI, Software Engineering, Data Science, and IT Operations, along with industry-leading external collaborators, to put the power of ML and computational tooling directly into researchers' day-to-day work.

Core ResponsibilitiesModel Deployment, Serving & Inference
  • Move trained models from research and experimentation into production, packaging, versioning, and promoting them across development, staging, and production environments and across cloud targets (AWS, Azure, GCP) and on-prem or hybrid infrastructure
  • Build and operate scalable inference services and APIs-batch, real-time, and streaming-delivering low-latency, high-throughput serving that meets researcher and downstream-system needs
  • Design and maintain model-serving infrastructure using containers and Kubernetes, with autoscaling, versioned rollouts (e.g., blue-green or canary), and rollback so updates ship without disrupting users
  • Integrate models into researcher-facing tools and enterprise systems, ensuring seamless interoperability and data flow across platforms
Data Pipelines & Readiness
  • Design, build, and maintain scalable, secure data pipelines-batch, change-data-capture (CDC), and streaming-that move and transform data across the platform, including the embedding, vectorization, and feature pipelines that feed downstream ML and LLM applications
  • Implement scalable storage and retrieval for large-scale structured and unstructured scientific data across cloud and on-prem or hybrid infrastructure
  • Build and operate automated data-readiness and quality-monitoring workflows for high-dimensional scientific and enterprise datasets, including multi-method anomaly and outlier detection across numerical and categorical data
  • Validate files for missing values, illegal characters, and structural issues, and build schema-drift detection with historical tracking and automated reporting-catching data-contract changes before they reach models and significantly reducing manual data QA
Model & Data Validation, Monitoring & Governance
  • Author, review, and validate model cards-verifying documented performance, intended use, limitations, data lineage, and evaluation results before models are promoted
  • Run and automate model validation and evaluation-reproducing metrics, checking calibration and performance against acceptance criteria, and gating promotion on the results
  • Implement production monitoring for model, data, and service health-latency, throughput, data and prediction drift, and quality-with alerting and proactive remediation
  • Define acceptance criteria, audit trails, and reproducible checks; adjudicate flagged data and model issues with data owners and scientists; and track and report operational metrics
Software & Platform Engineering
  • Design and develop robust, scalable, and secure software solutions with a hands-on approach, from architecture through implementation
  • Build and maintain microservices architectures and APIs (REST and GraphQL) that support model serving, data access, and tool-calling workflows
  • Implement infrastructure-as-code and CI/CD pipelines to automatically test and deploy model, data, and service updates, applying test-driven development to catch regressions early
  • Apply systems-engineering practices to distributed systems with high throughput and availability requirements, and troubleshoot complex issues across the model, data, and serving stack
Cross-functional Partnership
  • Collaborate within a team of engineers using best practices such as design reviews, code reviews, testing, and continuous integration and deployment
  • Partner with Lilly Research Labs, Data Science, AI/ML, and IT Operations to translate research and business requirements into technical solutions
  • Work with external, industry-leading collaborators to integrate models, data, and tooling into shared and federated workflows within Lilly's controlled cloud environment
  • Contribute to platform adoption through clear documentation, data dictionaries, runbooks, and support for internal end users
Required Qualifications
  • Ph.D. in Computer Science or a related computational field (e.g., Computational Science, Computational Biology, Bioinformatics, or a related quantitative computational discipline)
  • Hands-on experience in software engineering and architecture, with a proven track record of delivering complex, cross-functional solutions
  • Proficiency in a systems or object-oriented language (Go, Rust, Java, or C++) and a scripting language (Python and/or JavaScript)
  • Hands-on experience deploying to containers, serverless, Kubernetes, and other hosting targets
  • Experience deploying and serving machine learning models in production, including packaging, versioning, and promotion across environments
  • Experience building data pipelines and working with relational and non-relational data stores (e.g., PostgreSQL, MySQL, MongoDB)
  • Solid understanding of HTTP and RESTful APIs
  • Experience using CI tools to automatically test and CD tools to automatically deploy updates, and applying test-driven development to prevent feature regression
  • Experience applying systems-engineering concepts to distributed systems with high throughput and availability requirements
Preferred Qualifications
  • Experience integrating AI/ML models into production with a focus on scalability, performance, and reliability (MLOps)
  • Familiarity with MLOps and model-serving tooling (e.g., MLflow, Kubeflow, and model or artifact registries such as JFrog Artifactory)
  • Experience with model validation, evaluation, and model-card and documentation practices for model governance
  • Experience implementing data-quality, anomaly-detection, or schema-drift monitoring for production datasets
  • Familiarity with streaming and CDC tooling (e.g., Kafka, Kafka Streams, Spark Streaming) and big-data processing (Spark)
  • Familiarity with LLM application patterns-retrieval-augmented generation, tool-calling, and multi-agent orchestration-and with inference optimization
  • Experience with infrastructure-as-code (Terraform), service mesh, and cloud-native monitoring and observability
  • Exposure to drug discovery, life sciences, or healthcare data and workflows, including high-dimensional or biological datasets
  • Experience contributing to federated or collaborative ML and data initiatives across organizations

Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.


Lilly is proud to be an EEO Employer and does not discriminate on the basis of age, race, color, religion, gender identity, sex, gender expression, sexual orientation, genetic information, ancestry, national origin, protected veteran status, disability, or any other legally protected status.


Our employee resource groups (ERGs) offer strong support networks for their members and are open to all employees. Our current groups include: Africa, Middle East, Central Asia (AMECA), Black Employees at Lilly (BE@Lilly), Chinese Culture Network (CCN), EnAble, Evolve, Lilly Indian Network (LIN), Organization of Latinx at Lilly (OLA), Pride (LGBTQ+ Allies), Veterans Leadership Network (VLN) and Women's Initiative for Leading at Lilly (WILL).


Actual compensation will depend on a candidate's education, experience, skills, and geographic location. The anticipated wage for this position is

$151,500 - $244,200

Full-time equivalent employees also will be eligible for a company bonus (depending, in part, on company and individual performance). In addition, Lilly offers a comprehensive benefit program to eligible employees, including eligibility to participate in a company-sponsored 401(k); pension; vacation benefits; eligibility for medical, dental, vision and prescription drug benefits; flexible benefits (e.g., healthcare and/or dependent day care flexible spending accounts); life insurance and death benefits; certain time off and leave of absence benefits; and well-being benefits (e.g., employee assistance program, fitness benefits, and employee clubs and activities).Lilly reserves the right to amend, modify, or terminate its compensation and benefit programs in its sole discretion and Lilly's compensation practices and guidelines will apply regarding the details of any promotion or transfer of Lilly employees.

#WeAreLilly


What Eli Lilly and Company employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Eli Lilly logo

About Eli Lilly

Sourced by ZipRecruiter

Eli Lilly, based in Indianapolis, IN, US, is one of the pioneers in the pharmaceutical industry with a rich history dating back to 1876. This global pharmaceutical company focuses on discovering, developing, manufacturing and selling pharmaceutical products in approximately 120 countries. The company's product categories include endocrinology, oncology, cardiovascular, neuroscience, and immunology. Having invested over $9 billion in research and development in the past decade, Eli Lilly is also committed to creating high-quality medicines that meet real needs. As a recipient of several awards and recognitions, Eli Lilly is known for its focus on life-saving research and drug development. Their mission is to make medicines that help people live longer, healthier, and more active lives.

Industry

Pharmaceutical product wholesalers

Company size

10,000+ Employees

Headquarters location

Indianapolis, IN, US

Year founded

1876