1

Data Optimization Jobs in California (NOW HIRING)

Systems PhD - Software Engineer

Mountain View, CA · On-site

$205K - $244K/yr

Responsibilities : • Query compilation & optimization • Distributed query execution and scheduling • Vectorized engine execution • Data security • Resource Management • Transaction ...

Senior Application Data Engineer

Los Angeles, CA · On-site

$114K - $155K/yr

... optimization, ensure data integrity and accuracy. • Evaluate and make decisions regarding competing data design tradeoffs. • Advocate, defend and convert other engineering leaders to your own ...

Lead data architecture designs, drive data optimization, ensure data integrity and accuracy. * Evaluate and make decisions regarding competing data design tradeoffs. * Advocate, defend and convert ...

Sr. Data Engineer

Sacramento, CA · On-site

$124K - $149K/yr

Enterprise Analytics & Reporting, Cloud Migration & Server Administration, SAP Enterprise Support, Protocol & Data Optimization, Global Study & Contract Analytics Professional Certification * Tableau ...

AI/ML Architect

Los Angeles, CA · On-site

$68.75 - $88.25/hr

Data partitioning, file size tuning, and optimization strategies for large-scale pipelines. * Experience handling multi-terabyte structured time‑series workloads. * Ability to distill architectural ...

next page

Showing results 1-20

Data Optimization information

What are some typical challenges faced in a Data Optimization role?

Professionals in Data Optimization often encounter challenges such as working with incomplete or inconsistent datasets, integrating data from multiple sources, and ensuring data quality and accuracy throughout the optimization process. Balancing technical efficiency with business objectives and communicating complex analytical findings in easily understandable ways can also be demanding. Collaboration with cross-functional teams is frequent, requiring both strong technical and interpersonal skills. Overcoming these challenges helps ensure that optimization projects deliver meaningful value and measurable impact for the organization.

What is a Data Optimization job?

A Data Optimization job involves improving the efficiency, accuracy, and accessibility of data within an organization. Professionals in this role analyze large datasets, refine data structures, and implement strategies to enhance data processing and storage. They work with data engineers, analysts, and business teams to ensure data supports performance goals and decision-making. Common tasks include cleaning data, reducing redundancies, and optimizing database queries.

What are the key skills and qualifications needed to thrive in the Data Optimization position, and why are they important?

To thrive in Data Optimization, you need strong analytical skills, expertise in data modeling, and a solid foundation in statistics or mathematics, usually supported by a relevant degree. Familiarity with tools such as SQL, Python, R, and data visualization platforms like Tableau, as well as certifications in data analytics or optimization software, is highly beneficial. Effective communication, problem-solving abilities, and a collaborative mindset are key soft skills for this role. These competencies are crucial for translating complex data into actionable insights that drive business efficiency and performance improvements.

What are the most commonly searched types of Data Optimization jobs in California? The most popular types of Data Optimization jobs in California are:
What are popular job titles related to Data Optimization jobs in California? For Data Optimization jobs in California, the most frequently searched job titles are:
What job categories do people searching Data Optimization jobs in California look for? The top searched job categories for Data Optimization jobs in California are:
What cities in California are hiring for Data Optimization jobs? Cities in California with the most Data Optimization job openings:
Infographic showing various Data Optimization job openings in California as of July 2026, with employment types broken down into 57% Full Time, and 43% Contract. Highlights an 100% In-person job distribution.
Software Engineer

Software Engineer

Denken Solutions, Inc.

Foster City, CA

$100.70/hr

Contractor

Posted 9 days ago


Job description

On-Site role

Job Description:

  • Software Engineer, Autonomy Behavior ML Data Optimization Team.
  • Onsite in Foster City, CA | 5 days in office.
  • The Autonomy Behavior ML Data Optimization team is looking for a Software Engineer with strong data processing and pipeline engineering skills to build, scale, and optimize ScenarioScout, the scenario discovery platform at Client that empowers teams across the autonomous vehicle development lifecycle.
  • ScenarioScout transforms large-scale driving data into actionable insights by enabling fast, semantic similarity search over millions of driving scenario embeddings.
  • Client is building a future for Riders, not drivers.
  • At the heart of this vision, the Autonomy Behavior ML Data Optimization team develops models that forecast the behavior of all road agents around the Client robotaxi.
  • ScenarioScout is the useful data tool that enables teams to discover, curate, and validate driving scenarios from finding diverse training data for ML models to investigating safety-critical edge cases and enriching Focus Area (FA) datasets.
  • The core of this role is data pipeline engineering: designing and operating Airflow DAGs that orchestrate end-to-end dataset creation and refresh pipelines, building Ray-based distributed processing jobs for k-means clustering, embedding cache generation, FAISS index building, and dataset statistics computation over large Parquet datasets stored in S3.
  • The engineer will also work on the Python/FastAPI backend powering high-throughput embedding search and bulk operations, and contribute to the React.js frontend for UI-driven pipeline management, admin panel features, and search experience improvements.
  • The ideal candidate is someone who thrives in data-intensive environments and is comfortable working across the stack when needed, with particular depth in data processing and pipeline orchestration.
  • If you are excited about building data pipelines and processing systems that directly accelerate safe autonomous driving, enjoy working with large-scale distributed data infrastructure, and want to make a measurable impact on how a world-class AV company discovers and understands driving scenarios, this role is for you.

Responsibilities:

  • Build and maintain the Python/FastAPI backend powering high-throughput embedding search, bulk execution APIs, and dataset management endpoints.
  • Develop, maintain, and enhance dashboards and data visualization tools for data introspection and ad-hoc reporting.
  • Implement observability and monitoring (pipeline health, search hit rate, result download rate, system health metrics) to ensure platform reliability and data-driven iteration.
  • Design, build, and maintain Airflow DAGs that orchestrate end-to-end dataset creation and refresh pipelines, including embedding generation, distributed k-means clustering, FAISS index building, embedding cache construction, and dataset statistics computation.
  • Optimize data processing performance and pipeline reliability, including embedding generation pipeline improvements, cache building optimizations, and index construction tuning.
  • Enhance and operate scalable data processing pipelines for distributed k-means clustering, embedding cache building, FAISS index generation, and dataset statistics computation on large Parquet datasets stored in S3.
  • Own the ScenarioScout dataset lifecycle: full dataset creation, incremental refresh, bulk clustering assignment, dataset registry management, and embedding onboarding for new embedding types (e.g., non-QTP embeddings).
  • Collaborate with ML researchers on integrating new embedding types, improving embedding quality, and exploring LLM-powered natural language querying capabilities.
  • Contribute to the React.js frontend for admin panel features (dataset creation forms, pipeline monitoring, metadata field configuration), search experience improvements, and UI-driven pipeline management.
  • Deliver UI/UX improvements informed by user feedback, including visualization enhancements, visualization toggles, query result deduplication, shareable search result links, search interruption controls, and result grouping for overlapping scenarios.
  • Work closely with cross-functional stakeholders (Prediction, Data Optimization, Safety, QA) to understand requirements and translate them into scalable, maintainable software.


Qualifications:

  • 3+ years of professional software engineering experience with a focus on data processing and pipeline engineering.
  • Experience designing, building, and optimizing distributed data processing pipelines at scale (e.g., ETL/ELT) using technologies like Spark, Databricks, AWS EMR, AWS Batch, or Ray Core/Data.
  • Strong proficiency in Python with experience building production data pipelines and web services (FastAPI, Uvicorn, or similar async frameworks).
  • Experience building and maintaining data visualization dashboards, with proficiency in SQL, and familiarity with PySpark/Scala for large-scale data manipulation.
  • Experience with workflow orchestration tools (Airflow, Prefect, Dagster, or similar) for managing complex multi-step data processing pipelines.
  • Familiarity with vector similarity search and indexing technologies (FAISS, Annoy, ScaNN, Milvus, or similar).
  • Experience working with cloud infrastructure (AWS S3, EC2) and container orchestration (Kubernetes, Docker).
  • Strong understanding of data structures, algorithms, and performance optimization for data-intensive workloads.
  • Familiarity with React.js or a comparable modern frontend framework for building interactive data-driven applications.
  • Excellent communication and collaboration skills, ability to work effectively with ML researchers, data engineers, and product stakeholders.


Bonus Qualities:
Experience with any of the following:

  • Machine Learning concepts, particularly embeddings, clustering (k-means), and similarity search.
  • Building or operating ML serving infrastructure (e.g., TensorFlow Serving, TorchServe, Triton, or custom model serving like RayServe).
  • Large Language Model (LLM) integration for natural language interfaces or retrieval-augmented generation (RAG) workflows.
  • Building and operating production Airflow DAGs with complex task dependencies, retries, and monitoring.
  • Data catalog, dataset registry, or metadata management systems.
  • Visualization tools for robotics or autonomous driving data (RViz, Foxglove, or similar).
  • Observability and monitoring tooling (Prometheus, Grafana, Datadog, or similar).
  • CI/CD pipeline development and Bazel build systems.
  • Autonomous vehicle domain knowledge or experience with driving data (logs, scenarios, telemetry).
  • Full-stack development with experience owning applications end-to-end from frontend to infrastructure.
  • Building admin panels and dataset management UIs.

About us:
At our organization, we take our mission and values to heart! We are on a mission to offer more and better jobs all over the world! Our goal is to care for you while you care for our clients and get you paid the highest pay possible. All our associates working with us are expected to embrace our RACE values: R - Results Matter, A- Approachable, C - Care, and E - Emergency i.e. work with a sense of urgency.

For more relevant job opportunities please visit our website: Denken Solutions Careers