1

Synthetic Data Generation Jobs (NOW HIRING)

Working within a cross-functional team and reporting to a technical lead, you will operate across the machine learning development lifecycle, from data curation and synthetic data generation to model ...

Staying in sync with the latest state-of-the-art research in synthetic data generation and LLM training is key to success in this role. You will constantly lead original research initiatives through ...

Design synthetic task generation methods that produce diverse, realistic, and learnable tasks ... Experience with synthetic data research methods - please elaborate in your application * Strong ...

Data Scientist

Herndon, VA · On-site

$106K - $180K/yr

Working within a cross-functional team and reporting to a technical lead, you will operate across the machine learning development lifecycle, from data curation and synthetic data generation to model ...

Senior Robotics Data Engineer - Only W2

Warren, MI · On-site

$99K - $135K/yr

... and synthetic data generation. · Manage data versioning, metadata, and dataset governance to support model training, evaluation, and regression testing. · Collaborate with Robotics Perception ...

Showing results 21-40

Synthetic Data Generation information

See salary details

$31K

$93.2K

$169K

How much do synthetic data generation jobs pay per year?

As of Sep 6, 2026, the average yearly pay for synthetic data generation in the United States is $93,198.00, according to ZipRecruiter salary data. Most workers in this role earn between $54,500.00 and $144,500.00 per year, depending on experience, location, and employer.

What is synthetic data generation?

Synthetic data generation is the process of creating artificial datasets that mimic real-world data. This technique is used to supplement or replace actual data for purposes such as machine learning, software testing, and research, especially when real data is scarce, sensitive, or costly to obtain. Synthetic data can help improve model accuracy, protect privacy, and enable innovation by providing diverse and unbiased datasets. It is commonly used in fields like healthcare, finance, and autonomous vehicles.

What are the key skills and qualifications needed to thrive in synthetic data generation?

To excel in a Synthetic Data Generation role, you need a solid background in computer science, statistics, and data science, often supported by a relevant degree and experience in machine learning. Familiarity with tools such as Python, TensorFlow, PyTorch, and synthetic data generation platforms, as well as knowledge of privacy-preserving techniques, is typically required. Strong problem-solving abilities, creativity, and effective communication set top performers apart in this field. These skills and qualities are crucial for creating high-quality, realistic synthetic datasets that support robust AI model development while safeguarding sensitive information.

What are the main challenges faced by professionals working in synthetic data generation, and how can they be addressed?

Professionals in synthetic data generation often encounter challenges such as ensuring the generated data accurately represents real-world scenarios while maintaining privacy and data security. Balancing realism with anonymization is crucial, especially when synthetic data is used for AI model training or testing. Collaboration with data scientists, domain experts, and privacy officers is common to validate data utility and compliance with regulations. Staying current with advances in generative models and data validation techniques also helps address these challenges and contributes to career growth in this rapidly evolving field.

What is the difference between Synthetic Data Generation vs Data Analyst?

AspectSynthetic Data GenerationData Analyst
Required CredentialsKnowledge of data science, programming, and data privacyDegree in statistics, data science, or related field
Work EnvironmentData science teams, research labs, tech companiesBusiness environments, analytics teams, consulting firms
Industry UsageAI development, machine learning, data privacyBusiness insights, reporting, decision-making
Search & Comparison IntentUnderstanding data generation techniques, privacy solutionsAnalyzing data, generating reports, insights

While Synthetic Data Generation focuses on creating artificial data for privacy and model training, Data Analysts interpret existing data to provide business insights. Both roles require data-related skills but serve different purposes within the data ecosystem.

More about Synthetic Data Generation jobs

What cities are hiring for Synthetic Data Generation jobs?

Cities with the most Synthetic Data Generation job openings:

What states have the most Synthetic Data Generation jobs?

States with the most job openings for Synthetic Data Generation jobs include:

What job categories do people searching Synthetic Data Generation jobs look for?

The top searched job categories for Synthetic Data Generation jobs are:

Infographic showing various Synthetic Data Generation job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 85% Full Time, 11% Part Time, and 3% Contract. Highlights an 86% Physical, 3% Hybrid, and 11% Remote job distribution, with an average salary of $93,198 per year, or $44.8 per hour.

Senior Machine Learning Engineer - Physical AI and Synthetic Data Generation and Evaluation

Nvidia

Santa Clara, CA • On-site

$143K - $189K/yr

Full-time

This job post has expired 3 days ago. Applications are no longer accepted.


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 18 frontline employees who took The Breakroom Quiz

6th of 247 rated software companies


Job description

We are looking for outstanding Machine Learning Engineers to join our Physical AI teams. As the pioneers of the GPU-the visual cortex of modern computing-we are building the foundation for the next wave of AI that interacts with the physical world.

This role is at the forefront of Physical AI, developing sophisticated reasoning modules to build high-fidelity synthetic datasets. It leverages state-of-the-art multimodal models and diffusion techniques to simulate complex physical environments, ensuring our AI agents are trained on the most diverse and rigorous data possible. In particular, we will build advanced quality assurance technology to validate generated data outputs. We work closely with various users of synthetic datasets, including policy models. It extends an opportunity to contribute to the technology that will drive the cars of the future!

What you'll be doing:

  • Architect Generative Pipelines: Develop and implement advanced image and video generation/editing/reasoning models to produce high-fidelity synthetic data for Physical AI applications.

  • Multimodal Development: Build and fine-tune large-scale models, including VLMs, MLLMs, Generation models, applying transformer, auto-regressive and diffusion-based architectures. These models will take both visual and structured inputs (such as world model representations), and generate data and analyze consistency between scenarios intended by users and the generated data.

  • Controllable Synthesis: Apply and evolve user controls during data generation to ensure precise environmental and structural control over generated data.

  • Automated Quality Assurance for Sensor Data and Ego Policy: Build and test automated data QA pipeline using MLLMs and a mix of well known classical algorithms. In particular, build new capabilities to judge the quality of behavioral policies to ensure high quality data delivery for VLA.

  • Detailed Validation: Establish a strong mentality for KPI evaluation and validation to ensure the quality and physical accuracy of the synthetic releases. Establish a benchmark dataset. Design and validate KPI metric designs.

  • SOTA Data Engineering: Lead the generation of massive training datasets using various state-of-the-art tools and synthetic data mining techniques.

  • Contribute to the full lifecycle of ML software, including performance optimization, testing, and high-quality documentation.

What we need to see:

  • BS, MS, or PhD in Computer Science, Computer Graphics, Robotics, or a related field (or equivalent experience).

  • 12+ years of experience in ML software development.

  • Deep technical knowledge of image/video synthesis, including diffusion models and state-of-the-art multimodal methods.

  • Strong hands-on skills in major DNN libraries and computer languages including Python among others. Various hands on experience with workflow management and database to facilitate large scale training and data generation. Strong skills to optimize code efficiency is a huge plus.

  • Strong analytical and mathematical skills to bridge the gap between data-driven approaches and physical world constraints.

  • A collaborative outlook with outstanding communication skills, thriving in a tightly-knit team environment.

  • Experience in assessing the impact of synthetic data on model performance through metrics and systematic validation.

Ways to stand out from the crowd:

  • Experience with computer/GPU architecture to improve the performance during inference/training.

  • Familiarity with simulation platforms and deep understanding of 3D sensor modalities (Camera, Multi cameras, Lidar, Radar).

  • Experience with open source software.

We value different paths to technical excellence and welcome candidates who bring strong judgment, curiosity, and a collaborative approach. Come build the future of autonomous vehicle simulation with us!

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 224,000 USD - 356,500 USD for Level 5, and 272,000 USD - 431,250 USD for Level 6.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until August 1, 2026.

This posting is for an existing vacancy.

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US