1

Scale Ai Data Labeling Jobs (Flexible Options) in California

Experience scaling a multi-vendor or multi-region labeling operation. * Experience with golden set / ground truth construction. * Background in robotics, autonomous systems, LLM, or physical AI data.

Experience scaling a multi-vendor or multi-region labeling operation. * Experience with golden set / ground truth construction. * Background in robotics, autonomous systems, LLM, or physical AI data.

Showing results 21-40

Scale Ai Data Labeling information

What are popular job titles related to Scale Ai Data Labeling jobs in California? For Scale Ai Data Labeling jobs in California, the most frequently searched job titles are:
What job categories do people searching Scale Ai Data Labeling jobs in California look for? The top searched job categories for Scale Ai Data Labeling jobs in California are:
What cities in California are hiring for Scale Ai Data Labeling jobs? Cities in California with the most Scale Ai Data Labeling job openings:

Senior AI Data Pipeline Engineer (Autonomous Driving)

42dot

Sunnyvale, CA • On-site

$123K - $168K/yr

Full-time

Re-posted 5 days ago


Job description

Job Summary:
42dot is a mobility AI company committed to solving mobility challenges with software and AI. They are seeking a Senior AI Data Pipeline Engineer to build core data processing pipelines and datasets for autonomous driving algorithms, enhancing the efficiency of the ML model development lifecycle.
Responsibilities:
• Develop high scale, reliable data extraction pipeline to extract millions of raw data from data collection fleet and convert to high-value scene data
• Develop data labeling pipelines to perform the auto labeling inferences for autonomous driving algorithms
• Develop advanced autonomous driving data SDK, including scene data search, datasets preparation, dataset loading, etc.
• Build up the data lakehouse for autonomous driving scene dataset, including the sensor data, calibration data, as well as annotation data
• Dig into performance bottlenecks all along the data processing pipelines, from data processing latency, data search latency to Test Procedure (TP) coverage.
• Bootstrap and maintain infrastructure for data platform components—data processing pipeline, database, data lakehouse and data serving.
• Collaborate with cross-functional teams, including ML algorithm, ML application, and Cloud Infra to align data pipelines with overall autonomous driving system architecture.
Qualifications:
Required:
• Bachelor's degree or higher in Computer Science, Engineering, Robotics, or a similar technical field.
• Minimum of 7 years of experience in Data Engineering, DataOps or ML Platform roles
• Proficient in Python and solid experience in Python SDK development
• Solid hands-on experience with data pipeline job orchestration with Databricks Workflows or Apache Airflow, as well as integrating data pipelines with machine learning models
• Solid working experience in Databases (e.g., MongoDB, PostgreSQL, etc)
• Extensive experience with data technologies and architectures such as Data Warehouse (e.g., Hive) or Lakehouse (e.g., Delta Lake)
• Experience with Apache Spark or other big data computing engines
• Excellent leadership and communication skills, with a demonstrated ability to lead technical projects
Preferred:
• Experience with autonomous vehicle sensor data (e.g., LiDAR, camera, radar)
• Experience with ML model training lifecycle (e.g., data preparation, model training / validation / deployment, etc)
• Understanding of modern AI frameworks (e.g., PyTorch, TensorFlow etc.)
• Understanding data governance principles, data privacy regulations, and experience implementing security measures to protect data
Company:
42dot is a mobility AI company that specializes in autonomous transportation and mobility solutions. Founded in 2019, the company is headquartered in Seoul, KOR, with a team of 501-1000 employees. The company is currently Late Stage.