1

Data Ingestion Jobs (NOW HIRING)

Data Engineer, Scientific Data Ingestion

San Francisco, CA · On-site

$150K - $200K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

WHAT YOU WILL DO Build and own an AI-powered ingestion & normalization pipeline to import data from a wide variety of sources - unprocessed Excel/CSV uploads, lab and instrument exports, as well as ...

Data Engineer, Scientific Data Ingestion

San Francisco, CA · On-site

$150K - $200K/yr

  • Medical

  • Dental

  • Vision

  • Retirement

WHAT YOU WILL DO Build and own an AI-powered ingestion & normalization pipeline to import data from a wide variety of sources -- unprocessed Excel/CSV uploads, lab and instrument exports, as well as ...

Roku is seeking a Senior Solutions Engineer, Ad Data Ingestion to join our Solutions Engineering team supporting Roku's advertising business. Roku's advertising business runs on the quality of the ...

Sr. Solutions Engineer, Ad Data Ingestion

Boston, MA · On-site

$123K - $175K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Roku is seeking a Senior Solutions Engineer, Ad Data Ingestion to join our Solutions Engineering team supporting Roku's advertising business. Roku's advertising business runs on the quality of the ...

Sr. Solutions Engineer, Ad Data Ingestion

Austin, TX · On-site

$123K - $175K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Roku is seeking a Senior Solutions Engineer, Ad Data Ingestion to join our Solutions Engineering team supporting Roku's advertising business. Roku's advertising business runs on the quality of the ...

Sr. Solutions Engineer, Ad Data Ingestion

New York, NY · On-site

$123K - $175K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Roku is seeking a Senior Solutions Engineer, Ad Data Ingestion to join our Solutions Engineering team supporting Roku's advertising business. Roku's advertising business runs on the quality of the ...

Roku is seeking a Senior Solutions Engineer, Ad Data Ingestion to join our Solutions Engineering team supporting Roku's advertising business. Roku's advertising business runs on the quality of the ...

Sr. Solutions Engineer, Ad Data Ingestion

Boston, MA · On-site

$123K - $175K/yr

  • Medical

  • Life

  • PTO

Roku is seeking a Senior Solutions Engineer, Ad Data Ingestion to join our Solutions Engineering team supporting Roku's advertising business. Roku's advertising business runs on the quality of the ...

Sr. Solutions Engineer, Ad Data Ingestion

Santa Monica, CA · On-site

$123K - $175K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Roku is seeking a Senior Solutions Engineer, Ad Data Ingestion to join our Solutions Engineering team supporting Roku's advertising business. Roku's advertising business runs on the quality of the ...

next page

Showing results 1-20

Data Ingestion information

See salary details

$24.5K

$126.1K

$173K

How much do data ingestion jobs pay per year?

As of Aug 16, 2026, the average yearly pay for data ingestion in the United States is $126,123.00, according to ZipRecruiter salary data. Most workers in this role earn between $102,000.00 and $159,000.00 per year, depending on experience, location, and employer.

What are some common challenges faced in a data ingestion role, and how can they be addressed?

Professionals working in data ingestion often encounter challenges such as handling diverse data formats, ensuring data quality, and maintaining data pipeline reliability. Managing large volumes of data from multiple sources can lead to issues with consistency, latency, and error handling. To address these, it is essential to implement robust validation processes, use scalable data ingestion tools, and closely collaborate with data engineering and source system teams. Continuous monitoring and regular pipeline optimization are also key practices to ensure smooth data flow and integrity.

What is data ingestion?

Data ingestion is the process of collecting and importing data from various sources into a storage or processing system, such as a database, data warehouse, or data lake. It involves transferring raw data, which may come from different formats and platforms, into a centralized location for analysis and further processing. Data ingestion can be done in real-time (streaming) or in batches, depending on business needs and data velocity. This process is crucial for organizations to make informed decisions based on comprehensive, up-to-date information.

What is the difference between Data Ingestion vs Data Analyst?

AspectData IngestionData Analyst
Required CredentialsKnowledge of ETL tools, SQL, basic scriptingDegree in statistics, data science, or related field; SQL, Excel, visualization skills
Work EnvironmentData engineering teams, IT departments, cloud platformsBusiness units, analytics teams, reporting environments
Industry UsageData pipelines, data warehouses, big data platformsData interpretation, reporting, decision support

Data ingestion involves collecting and importing data into storage systems, focusing on data pipelines and infrastructure. Data analysts interpret and analyze this data to generate insights. While data ingestion prepares data for analysis, data analysts focus on understanding and communicating data findings. Both roles are essential in data-driven organizations but serve different functions within the data lifecycle.

What are the key skills and qualifications needed to thrive as a data ingestion specialist?

To thrive as a Data Ingestion Specialist, you need a strong understanding of data pipelines, ETL processes, and experience with databases, often supported by a degree in computer science or related fields. Familiarity with tools such as Apache Kafka, Apache NiFi, SQL, and cloud platforms like AWS or Azure is typically required, along with knowledge of data integration frameworks. Strong problem-solving skills, attention to detail, and effective communication set top performers apart in this role. These skills are crucial to ensure the accurate, efficient, and secure movement of data across systems, supporting reliable analytics and business decisions.
More about Data Ingestion jobs

What cities are hiring for Data Ingestion jobs?

Cities with the most Data Ingestion job openings:

What states have the most Data Ingestion jobs?

States with the most job openings for Data Ingestion jobs include:

What job categories do people searching Data Ingestion jobs look for?

The top searched job categories for Data Ingestion jobs are:

Infographic showing various Data Ingestion job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 84% Full Time, 11% Part Time, and 4% Contract. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution, with an average salary of $126,123 per year, or $60.6 per hour.

Data Engineer, Scientific Data Ingestion

Mithrl

San Francisco, CA • On-site

Full-time

Re-posted 21 days ago


Job description

Job Summary:
Mithrl is a pioneering company focused on accelerating the delivery of novel drugs and therapies through AI technology. They are seeking a Data Engineer to build and manage an AI-powered data ingestion and normalization pipeline, ensuring high-quality data is prepared for analysis and integration into their systems.
Responsibilities:
• Build and own an AI-powered ingestion & normalization pipeline to import data from a wide variety of sources — unprocessed Excel/CSV uploads, lab and instrument exports, as well as processed data from internal pipelines.
• Develop robust schema mapping, coercion, and conversion logic (think: units normalization, metadata standardization, variable-name harmonization, vendor-instrument quirks, plate-reader formats, reference-genome or annotation updates, batch-effect correction, etc.).
• Use LLM-driven and classical data-engineering tools to structure “semi-structured” or messy tabular data — extracting metadata, inferring column roles/types, cleaning free-text headers, fixing inconsistencies, and preparing final clean datasets.
• Ensure all transformations that should only happen once (normalization, coercion, batch-correction) execute during ingestion — so downstream analytics / the AI “Co-Scientist” always works with clean, canonical data.
• Build validation, verification, and quality-control layers to catch ambiguous, inconsistent, or corrupt data before it enters the platform.
• Collaborate with product teams, data science / bioinformatics colleagues, and infrastructure engineers to define and enforce data standards, and ensure pipeline outputs integrate cleanly into downstream analysis and storage systems.
Qualifications:
Required:
• 5+ years of experience in data engineering / data wrangling with real-world tabular or semi-structured data.
• Strong fluency in Python, and data processing tools (Pandas, Polars, PyArrow, or similar).
• Excellent experience dealing with messy Excel / CSV / spreadsheet-style data — inconsistent headers, multiple sheets, mixed formats, free-text fields — and normalizing it into clean structures.
• Comfort designing and maintaining robust ETL/ELT pipelines, ideally for scientific or lab-derived data.
• Ability to combine classical data engineering with LLM-powered data normalization / metadata extraction / cleaning.
• Strong desire and ability to own the ingestion & normalization layer end-to-end — from raw upload → final clean dataset — with an eye for maintainability, reproducibility, and scalability.
• Good communication skills; able to collaborate across teams (product, bioinformatics, infra) and translate real-world messy data problems into robust engineering solutions.
Preferred:
• Familiarity with scientific data types and “modalities” (e.g. plate-readers, genomics metadata, time-series, batch-info, instrumentation outputs).
• Experience with workflow orchestration tools (e.g. Nextflow, Prefect, Airflow, Dagster), or building pipeline abstractions.
• Experience with cloud infrastructure and data storage (AWS S3, data lakes/warehouses, database schemas) to support multi-tenant ingestion.
• Past exposure to LLM-based data transformation or cleansing agents — building or integrating tools that clean or structure messy data automatically.
• Any background in computational biology / lab-data / bioinformatics is a bonus — though not required.
Company:
Mithrl is a software development company that builds the custom workflows for NGS data on-demand. Founded in 2023, the company is headquartered in San Francisco, USA, with a team of 51-200 employees. The company is currently Early Stage.