1

Data Preprocessing Jobs in Hercules, CA (NOW HIRING)

Experiencewith data preprocessing, feature engineering, and model evaluationtechniques. * Knowledgeof deep learning architectures (CNNs, RNNs, Transformers) and theirapplications. * Proficiencyin ...

Build and maintain data preprocessing and data generation pipelines to support model training and evaluation. * Run training and fine-tuning workflows end-to-end and iterate quickly on performance ...

AI/ML Engineer

San Francisco, CA · On-site

$140 - $200/hr

Experience with data preprocessing and feature engineering * PhD or Master's in Computer Science, AI, or related field (Nice to have) * Experience with large language models (GPT, BERT, etc.) (Nice ...

... data analysis, including preprocessing, feature engineering, and leveraging Generative AI algorithms for novel solutions. • Lead cross-functional collaborations to integrate Generative AI models ...

Work with large-scale datasets and utilize data preprocessing techniques to ensure high-quality input for training and production * Implement and maintain efficient data storage and retrieval ...

... data analysis, including preprocessing, feature engineering, and leveraging Generative AI algorithms for novel solutions. · Lead cross-functional collaborations to integrate Generative AI models ...

next page

Showing results 1-20

Data Preprocessing information

See Hercules, CA salary details

$50.8K

$182.2K

$268.9K

How much do data preprocessing jobs pay per year?

As of Aug 22, 2026, the average yearly pay for data preprocessing in Hercules, CA is $182,199.00, according to ZipRecruiter salary data. Most workers in this role earn between $147,400.00 and $187,700.00 per year, depending on experience, location, and employer.

What is data preprocessing?

Data preprocessing is the process of cleaning, transforming, and organizing raw data into a usable format for analysis or machine learning. It involves steps such as handling missing values, removing duplicates, normalizing or scaling data, and encoding categorical variables. Proper data preprocessing helps improve the quality and performance of predictive models by ensuring the data is accurate, consistent, and suitable for analysis.

What are the key skills and qualifications needed to thrive as a data preprocessing specialist, and why are they important?

To thrive as a Data Preprocessing Specialist, you need a strong background in statistics, data cleaning, and data transformation, often supported by a degree in computer science, data science, or a related field. Proficiency with tools such as Python (pandas, NumPy), SQL, and data visualization platforms is typically essential, along with familiarity with data management systems. Attention to detail, problem-solving abilities, and effective communication are standout soft skills in this position. These skills are crucial for ensuring high-quality, reliable datasets that underpin accurate data analysis and machine learning outcomes.

What are some common challenges faced in a data preprocessing role, and how can they be effectively managed?

Professionals in Data Preprocessing often encounter challenges such as handling incomplete or inconsistent data, managing large datasets, and ensuring data quality before analysis. Addressing these issues typically involves using specialized tools to automate data cleaning, establishing clear data validation rules, and collaborating closely with data engineers and analysts. Staying updated with best practices and leveraging scripting languages like Python or R can also streamline the preprocessing workflow, making it easier to deliver reliable and accurate datasets for downstream analysis.

What is the difference between Data Preprocessing vs Data Analysis?

AspectData PreprocessingData Analysis
Primary FocusCleaning, transforming, and preparing raw data for analysisInterpreting data to extract insights and support decision-making
Skills RequiredData cleaning, scripting, understanding of data formatsStatistical analysis, data visualization, critical thinking
Work EnvironmentData engineering teams, data science projectsBusiness intelligence, research, data science teams
Tools UsedPython, R, SQL, ETL toolsExcel, Tableau, R, Python, statistical software

While data preprocessing involves preparing raw data for analysis by cleaning and transforming it, data analysis focuses on interpreting the prepared data to uncover trends and insights. Both roles are essential in the data pipeline but serve different purposes in the data lifecycle.

What are popular job titles related to Data Preprocessing jobs in Hercules, CA?

For Data Preprocessing jobs in Hercules, CA, the most frequently searched job titles are:

What job categories do people searching Data Preprocessing jobs in Hercules, CA look for?

The top searched job categories for Data Preprocessing jobs in Hercules, CA are:

What cities near Hercules, CA are hiring for Data Preprocessing jobs?

Cities near Hercules, CA with the most Data Preprocessing job openings:

Infographic showing various Data Preprocessing job openings in Hercules, CA as of August 2026, with employment types broken down into 1% As Needed, 80% Full Time, 16% Part Time, and 3% Contract. Highlights an 87% Physical, 3% Hybrid, and 10% Remote job distribution, with an average salary of $182,199 per year, or $87.6 per hour.

Research Scientist - Vision Data Infrastructure (San Francisco)

Storm3

San Francisco, CA • On-site

$250K - $600K/yr

Full-time

Medical, Dental, Vision

This job post has expired today. Applications are no longer accepted.


Job description

Research Scientist - Vision Data Infrastructure

This role is offered by Storm3. Your actual pay will be based on your skills and experience — talk with your recruiter for details.

Base pay range $250,000 – $600,000 per year.

⚡ Research Scientists/Engineers (all levels)

🔍 Focus on Vision Data Infrastructure

Come join one of the only research institutions globally providing the resources to compete with top AI companies — tens of thousands of GPUs to explore state‑of‑the‑art research in LLMs, multimodal and agentic AI.

We are seeking AI talent with expertise in building scalable pipelines for vision data to support image/video generative training and multimodal alignment. You’ll design high-performance pipelines for large-scale image and video datasets, enabling efficient pretraining, alignment, and simulation-based data generation.

Responsibilities:

  • Vision Data Sourcing & Curation
    • Collect and organize image and video data from open datasets and the web.
    • Handle data cleaning, filtering, deduplication, and metadata generation.
    • Ensure ethical and compliant data collection at scale.
  • Processing & Augmentation
    • Build high-throughput pipelines for vision data preprocessing (frame extraction, resolution normalization, format conversion, latent caching).
    • Implement GPU-accelerated augmentation and distributed data loading (WebDataset, TFRecords, Parquet).
  • Synthetic & Simulation‑Based Data Generation
    • Use simulation tools (Unreal Engine 5, Isaac Sim, Unity) to generate high-quality synthetic vision data.
    • Create specialized datasets for VLM training, visual reasoning, and agent interaction.

Requirements:

  • Strong experience with data engineering, computer vision, or machine learning infrastructure.
  • Expertise in building and scaling ETL/data pipelines for large unstructured datasets.
  • Proficiency with Python, PyTorch, and distributed data frameworks (Ray, Spark, Dask).
  • Experience with WebDataset, TFRecords, Parquet, or similar high-throughput data formats.
  • Familiarity with GPU‑accelerated preprocessing, NVIDIA DALI, or equivalent systems.
  • Understanding of image/video codecs, data compression, and cloud storage optimization.

Preferred Experience:

  • Prior work with simulation-based or synthetic data generation using Unreal Engine, Isaac Sim, or Unity.
  • Experience curating datasets for multimodal or vision‑language model training.
  • Knowledge of data ethics, privacy, and compliance frameworks for large-scale AI datasets.
  • Experience contributing to open datasets or data-centric AI research.

Why apply:

  • Opportunity to join a fast-growing core team that is already pushing AI breakthroughs.
  • Highly competitive salary package.
  • Work alongside ambitious and bright superstars from tech and academia.
  • Medical, Dental and Vision Insurance.

Interested? Please contact stefani.lukic@storm3.com.

Seniority level: Mid‑Senior level

Employment type: Full‑time

Job function: Research and Engineering

Industries: Software Development and Research Services

#J-18808-Ljbffr