Software Engineer - ML Data Infrastructure & Pipeline Engineering
Scope: Contract
Duration: 6 Months (Potential Extension)
Hours: 40 Hours/Week
Location: Foster City, CA
Onsite Requirement: 5 Days/Week Preferred (Hybrid flexibility may be considered for exceptional candidates)
Pay Range: $90-$95.79/hour (W2)
Ideal Candidate
A mid-level software engineer with strong C++ skills and experience processing large-scale datasets. Candidates who have built robust data pipelines, backend services, or distributed systems (and have exposure to embeddings, vector search, similarity search, or ML infrastructure) will be especially successful in this role.
Interview Process
Round 1: C++ Coding Assessment
Round 2: Systems Design / Technical Interview
Final Round: Hiring Team Interview
Overview
We are partnering with a leading autonomous vehicle company seeking a Software Engineer to help build and scale the data infrastructure that supports production machine learning and autonomous vehicle systems.
This team is responsible for preparing and processing large-scale datasets used to train production ML models. The engineer will focus on improving the reliability, scalability, automation, and performance of the pipelines and services that power data preparation workflows.
The ideal candidate has experience building robust data processing systems, working with large volumes of data, and developing production-grade software in distributed environments. Experience with vector search, embeddings, similarity search, or ML infrastructure is highly desirable.
What You'll Do
Data Processing & Pipeline Engineering
- Design, build, and maintain large-scale data processing pipelines
- Improve the stability, reliability, and automation of existing data workflows
- Process and manage massive datasets in distributed computing environments
- Optimize performance and scalability across data infrastructure
- Build systems that support continuous, production-grade data processing
Backend & Infrastructure Development
- Develop and maintain backend services that support data preparation workflows
- Build infrastructure that enables efficient dataset generation and management
- Create robust, maintainable software focused on operational excellence
- Partner with engineering teams to improve service reliability and system performance
Machine Learning Data Infrastructure
- Support workflows used to prepare training datasets for production ML models
- Contribute to infrastructure that enables embedding generation, vector processing, and similarity search
- Help scale systems that support ML and LLM-related development efforts
Required Qualifications
- 3+ years of software engineering experience
- Strong C++ development skills (required)
- Experience building and maintaining large-scale data processing pipelines
- Experience working with high-volume datasets in distributed or scalable environments
- Strong understanding of data processing, data infrastructure, and system reliability
- Experience building production-quality backend services or infrastructure
- Familiarity with SQL and large-scale data operations
- Ability to write efficient, scalable, and maintainable code
Technologies & Backgrounds That Translate Well
Candidates may come from backgrounds involving:
- Distributed Data Processing
- Data Infrastructure Engineering
- ML Platform Engineering
- Search Infrastructure
- Backend Systems Engineering
- Large-Scale Analytics Platforms
Relevant technologies may include:
- C++
- Spark
- Databricks
- Presto
- Distributed Compute Platforms
- Vector Databases
- Similarity Search Systems
- Embedding Pipelines
We are looking forward to your application! No remote candidates will be considered; please only apply if you are able to be on site in Foster City, CA.