1

Machine Learning Data Linguist Jobs (NOW HIRING)

The Fraud & Machine Learning team is the secret sauce behind Extend's post-purchase protection platform. As a Senior ML Data Scientist, you will own the development of cutting-edge machine learning ...

The Fraud & Machine Learning team is the secret sauce behind Extend's post-purchase protection platform. As a Senior ML Data Scientist, you will own the development of cutting-edge machine learning ...

We are seeking an experienced Machine Learning Data Engineer to develop, operationalize, and continuously improve production bioinformatics pipelines and data products that support AI-driven ...

We are looking for a summer intern to join our machine learning team. This team focuses on the ... Enrolled in a data science, computer science, mathematics, linguistic, or educational measurement ...

$11.50 - $15.50/hr

Machine Learning Data Associate based in Netherlands. This is an opportunity to lead engaging, high-impact virtual learning experiences for diverse audiences developing skills in machine learning and ...

$90 - $200/hr

CapTech Machine Learning Engineers are responsible for designing and implementing data-driven solutions for our clients, with a specific focus on building and deploying scalable machine learning ...

Coordinate data collection and annotation efforts. * Work with real-time data and content coming from various data sources. * Manage machine learning data pipelines. * Design tests for machine ...

Coordinate data collection and annotation efforts. * Work with real-time data and content coming from various data sources. * Manage machine learning data pipelines. * Design tests for machine ...

Showing results 41-60

Machine Learning Data Linguist information

See salary details

$52K

$73.4K

$95.5K

How much do machine learning data linguist jobs pay per year?

As of Aug 23, 2026, the average yearly pay for machine learning data linguist in the United States is $73,373.00, according to ZipRecruiter salary data. Most workers in this role earn between $66,500.00 and $79,000.00 per year, depending on experience, location, and employer.

What is a machine learning data linguist?

A Machine Learning Data Linguist is a specialist who works at the intersection of linguistics and artificial intelligence. They are responsible for annotating, curating, and analyzing language data to train and improve machine learning models, especially those focused on natural language processing (NLP). Their work often includes tasks like labeling text, refining speech recognition data, and ensuring that language models understand context, grammar, and cultural nuances. This role is essential in developing accurate and inclusive AI systems that interact with human language.

How does a machine learning data linguist typically collaborate with engineers and data scientists on projects?

A Machine Learning Data Linguist works closely with engineers and data scientists by providing linguistic insights and ensuring that language data is accurately annotated and interpreted. They often participate in cross-functional meetings to define project goals, clarify annotation guidelines, and review model outputs for linguistic quality. This collaboration helps bridge the gap between technical development and language-specific nuances, leading to more effective and culturally accurate machine learning models. Effective communication and a strong understanding of both linguistic theory and technical requirements are vital in this collaborative environment.

What are the key skills and qualifications needed to thrive as a machine learning data linguist, and why are they important?

To thrive as a Machine Learning Data Linguist, you need expertise in linguistics, data annotation, and a strong understanding of language structures, often supported by a degree in linguistics or computational linguistics. Familiarity with annotation tools, data labeling platforms, and programming languages like Python is typically required. Strong attention to detail, analytical thinking, and clear communication are essential soft skills for accurately interpreting and conveying linguistic phenomena. These skills ensure high-quality language data, which is critical for developing effective and unbiased machine learning models.
More about Machine Learning Data Linguist jobs

What cities are hiring for Machine Learning Data Linguist jobs?

Cities with the most Machine Learning Data Linguist job openings:

What states have the most Machine Learning Data Linguist jobs?

States with the most job openings for Machine Learning Data Linguist jobs include:

What job categories do people searching Machine Learning Data Linguist jobs look for?

The top searched job categories for Machine Learning Data Linguist jobs are:

Infographic showing various Machine Learning Data Linguist job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 85% Full Time, 11% Part Time, and 3% Contract. Highlights an 85% Physical, 4% Hybrid, and 11% Remote job distribution, with an average salary of $73,373 per year, or $35.3 per hour.

Machine Learning Data Engineer (DataOps), Materra

X Development, LLC

Mountain View, CA โ€ข On-site

$166 - $244/hr

Other

Posted 18 days ago


Job description

Software Engineering Mountain View, CA About the team

Materra is on a mission to radically reduce global waste and move to a true circular economy. The team has developed technology that identifies waste material at the molecular levelโ€”starting with plastics. Materra works with industry partners to improve the way recycling centers process plastics using AI and robotics, to make recycling more affordable and scalable.

About the Role

We are looking for a Machine Learning Data Engineer (DataOps) to build and unify the data infrastructure that powers our model training pipelines. In this role, you will lead the effort to consolidate fragmented data sources into a cohesive, high-quality data foundation.

Your primary focus will be designing automated ingestion pipelines, establishing data quality validation frameworks, and managing dataset versioning to support our machine learning training loops. You will bridge the gap between operations, remote annotation teams, and machine learning engineers to ensure our models are trained on reliable, well-structured data.

Key Responsibilities
  • Architect and build automated ETL (Extract, Transform, Load) and ELT (Extract, Load, Transform) data pipelines to aggregate, clean, and harmonize data from disparate sources, databases, and operational ingestion flows.
  • Implement DataOps practices, including data quality monitoring, automated schema validation, and anomaly detection to catch corrupt or mislabeled data early.
  • Standardize and integrate third-party annotation workflows and remote labeling feeds into unified datasets ready for model training.
  • Design and maintain dataset versioning and storage systems to allow reproducible machine learning experiments and seamless data retrieval.
  • Collaborate with machine learning engineers and operations teams to translate raw material, form factor, and sensor metadata into structured training features.
Requirements
  • Education: Degree in Computer Science, Data Engineering, Software Engineering, or a related technical field.
  • Data Engineering & Architecture: 3+ years experience building scalable data pipelines, managing relational and non-relational databases, and unifying fragmented data storage systems.
  • Modern Python Proficiency: Expertise in Python and data manipulation libraries (e.g., Pandas, NumPy, or SQL).
  • Data Quality & DataOps: Practical experience implementing automated data validation, quality control frameworks, and dataset versioning practices.
  • ML Data Lifecycle Understanding: Hands-on experience structuring datasets specifically for machine learning workflows, including handling annotations, metadata tracking, and training set curation.
Preferred Skills
  • Google Cloud Ecosystem: Hands-on experience with Google Cloud platform tools (e.g., BigQuery, Cloud Storage, Dataflow, Dataproc, Vertex AI Data Pipelines).
  • Workflow Orchestration: Experience managing pipelines using Google Cloud Composer or equivalent orchestration frameworks (e.g., Apache Airflow, Prefect, Dagster).
  • Multimodal / Unstructured Data: Experience handling mixed data types, including image datasets, sensor metadata, and unstructured physical property records.
  • Annotation Platform Integration: Familiarity with data labeling platforms, human-in-the-loop workflows, or integrating third-party annotation APIs.
  • Validation & Versioning Tooling: Exposure to data quality and ML versioning tools (e.g., Great Expectations, DVC, or TFX/Data Validation).

The US base salary range for this full-time position is $166,000 - $244,000 + bonus + equity + benefits. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your location during the hiring process.

Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, equity, or benefits.

An Equal Opportunity Workplace

At X, we don't just accept difference - we celebrate it, we support it, and we thrive on it for the benefit of our employees, our products and our community. We are proud to be an equal opportunity workplace and is an affirmative action employer. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or Veteran status. We also consider qualified applicants regardless of criminal histories, consistent with legal requirements.

#J-18808-Ljbffr