1

Data Engineer Ml Jobs in California (NOW HIRING)

As a Software Engineer on the Machine Learning Data Platform team at Liftoff, you will: * Work with an experienced team of ML, Software, and Infrastructure Engineers that are building the ML platform ...

Software Engineer, ML Data

San Francisco, CA · On-site

$180K - $230K/yr

  • Medical

  • Dental

  • Vision

As a Software Engineer on the Machine Learning Data Platform team at Liftoff, you will: * Work with an experienced team of ML, Software, and Infrastructure Engineers that are building the ML platform ...

Data Engineer

El Segundo, CA · On-site

$122K - $146K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

About the Role As a Data Engineer at Circadia Health, you will play a critical role in building and ... Reporting directly to the CTO, you will work closely with backend engineers, ML engineers, clinical ...

Data Engineer

El Segundo, CA

$122K - $146K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

About the Role As a Data Engineer at Circadia Health, you will play a critical role in building and ... Reporting directly to the CTO, you will work closely with backend engineers, ML engineers, clinical ...

About the Role As a Data Engineer at Circadia Health, you will play a critical role in building and ... Reporting directly to the CTO, you will work closely with backend engineers, ML engineers, clinical ...

Data Engineer

Sunnyvale, CA · On-site

$136K - $163K/yr

AI ML Data Engineer Location : Sunnyvale, CA (3 days work from office) Minimum 6 to 12 years of experience as data engineer in AI ML. * Snowflake and Python/Scala/Java * SQL, No SQL database, Hadoop ...

Staff ML Data Engineer (Datagrid)

San Francisco, CA · On-site

$134K - $162K/yr

We're looking for a Staff ML Data Engineer to join Procore's AI & Frontier Models organization. In this role, you'll be responsible for designing and building the data systems that power frontier ...

Staff ML Data Engineer (Datagrid)

San Francisco, CA · On-site

$134K - $162K/yr

We're looking for a Staff ML Data Engineer to join Procore's AI & Frontier Models organization. In this role, you'll be responsible for designing and building the data systems that power ...

Senior ML Data Engineer, MLO

Cupertino, CA · On-site

$68.75 - $91/hr

Minimum Qualifications 7+ years of industry experience as a data engineer serving various ML applications (vision domain preferred) Bachelor's degree in Computer Science or related field Preferred ...

next page

Showing results 1-20

Data Engineer Ml information

What are the key skills and qualifications needed to thrive as a data engineer ML?

To thrive as a Data Engineer ML, you need strong programming skills (especially in Python or Scala), knowledge of data modeling, and a solid foundation in database technologies, typically supported by a degree in computer science or a related field. Familiarity with big data frameworks (like Spark or Hadoop), cloud platforms (AWS, GCP, or Azure), and ETL tools, as well as relevant certifications, is highly beneficial. Excellent problem-solving abilities, teamwork, and clear communication help you collaborate with data scientists and stakeholders effectively. These skills are essential for building robust data pipelines and infrastructure that enable scalable, high-quality machine learning solutions.

What does a data engineer ML do?

A Data Engineer ML (Machine Learning) is responsible for designing, building, and maintaining the data pipelines and infrastructure necessary for machine learning applications. They clean, process, and organize large datasets to ensure data quality and accessibility for data scientists and ML engineers. In addition, they may work on deploying machine learning models to production environments and optimizing data workflows for efficiency and scalability.

What is the difference between Data Engineer Ml vs Data Scientist?

AspectData Engineer MlData Scientist
Required CredentialsBachelor's in CS, Data Engineering certificationsBachelor's/Master's in CS, Data Science certifications
Work EnvironmentBuilding data pipelines, managing databasesAnalyzing data, creating models
Employer & Industry UsageTech companies, finance, healthcareResearch institutions, tech firms, finance

Data Engineer Ml focuses on developing and maintaining data infrastructure and pipelines, while Data Scientists analyze data and build predictive models. Both roles often collaborate but serve different functions within data teams.

How do data engineer ML roles typically collaborate with data scientists and machine learning engineers on projects?

Data Engineer ML professionals work closely with data scientists and machine learning engineers by building and maintaining robust data pipelines, ensuring clean and reliable datasets are readily available for modeling and analysis. They often participate in meetings to understand model requirements, help optimize data storage for performance, and support the deployment of machine learning models into production environments. Effective collaboration involves continuous communication to troubleshoot data issues, implement data validation, and scale solutions as project needs evolve. This teamwork ensures that data-driven projects move efficiently from experimentation to deployment.
What job categories do people searching Data Engineer Ml jobs in California look for? The top searched job categories for Data Engineer Ml jobs in California are:
What cities in California are hiring for Data Engineer Ml jobs? Cities in California with the most Data Engineer Ml job openings:
Infographic showing various Data Engineer Ml job openings in California as of August 2026, with employment types broken down into 1% As Needed, 82% Full Time, 12% Part Time, 2% Temporary, and 3% Contract. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution.

Staff+ Data Engineer (ML Infrastructure)

Sanas

Palo Alto, CA • On-site

Full-time

Re-posted 6 days ago


Job description

Job Summary:
Sanas is pioneering the future of human communication with its innovative speech AI platform. The Staff Data Engineer will own the infrastructure for processing raw audio data, ensuring high-quality training-ready data for AI models, while collaborating with AI research scientists and ML engineers.
Responsibilities:
• Design and implement large-scale data pipelines that ingest, transform, validate, and serve high-quality audio and metadata for AI model training, evaluation, and product telemetry.
• Own the lakehouse architecture — table format choices (Iceberg vs. Delta Lake), partitioning strategies, metadata management, and schema evolution — with a bias toward reproducibility and auditability.
• Build and maintain batch and streaming pipelines using Spark, Flink, and orchestration tooling (Airflow or Dagster), with a clear-eyed view of when each is the right tool.
• Develop and maintain pipelines purpose-built for the unique challenges of audio data: large file volumes, time-series feature extraction, speaker and language metadata, and annotation versioning.
• Build tooling that supports the full audio data lifecycle — from raw ingestion and quality filtering through augmentation, segmentation, and training split generation — with reproducibility guarantees at every stage.
• Partner with ML engineers and research scientists to design data schemas, sampling strategies, and evaluation datasets that accurately reflect production conditions.
• Own data pipelines that feed human-in-the-loop annotation workflows — ensuring clean round-trips between raw data, labeling platforms, and training-ready outputs.
• Instrument pipelines with observability, data quality checks, lineage tracking, and alerting — so failures surface fast and root causes are traceable.
• Drive build vs. buy decisions for data quality, observability, and cataloging tooling with a clear framework grounded in Sanas's scale and roadmap.
• Own disaster recovery design for critical data assets — training datasets, evaluation benchmarks, and model checkpoints.
• Set the technical bar for the data engineering team — review designs and code, establish patterns, and document decisions in a way that raises the floor for everyone.
• Work cross-functionally with AI research, infrastructure, product, and legal to align data architecture with business needs and regulatory requirements.
• Contribute to hiring — identify strong candidates, conduct technical interviews, and help define what great looks like for data engineering at Sanas.
Qualifications:
Required:
• 5+ years of experience in data engineering, ML infrastructure, or data platform roles.
• Deep expertise building distributed batch and streaming data systems in production.
• Strong command of data processing frameworks: Spark, Flink, and Ray; and orchestrators: Airflow or Dagster.
• Hands-on experience with cloud data platforms — Snowflake, Databricks, or ClickHouse — and object storage (S3, GCS) on AWS or GCP.
• Solid understanding of data lifecycle management: privacy, security, compliance, and reproducibility from ingestion through model training.
• Proven ability to work directly with ML researchers and engineers to translate model requirements into data infrastructure decisions.
Preferred:
• Direct experience with audio data pipelines — file handling at scale, time-series features, speaker metadata, or audio annotation tooling.
• Familiarity with ASR, TTS, or speech enhancement model training workflows and the data requirements specific to each.
• Experience with MLOps tooling — experiment tracking, dataset versioning (DVC, LakeFS), and training pipeline orchestration.
Company:
Sanas is a real-time speech-understanding platform that modulates accents while preserving voices and emotions for natural interactions. Founded in 2020, the company is headquartered in Palo Alto, USA, with a team of 51-200 employees. The company is currently Growth Stage.