1

Synthetic Data Generation Jobs (NOW HIRING)

Build synthetic data generation pipelines using LLM-based methods and automated quality evaluation, producing datasets that improve the pre- and post-training of LLMs such as Nemotron - reasoning ...

Design synthetic task generation methods that produce diverse, realistic, and learnable outputs ... Hands-on experience applying synthetic data research methods to build end-to-end data generation ...

Design synthetic task generation methods that produce diverse, realistic, and learnable outputs ... Hands-on experience applying synthetic data research methods to build end-to-end data generation ...

Data Engineer

Buffalo, NY · On-site

$70 - $75/hr

Synthetic data generation * Experience with cloud platforms (Azure strongly preferred) * Ability to work with and across: * Legacy systems * Modern cloud environments * Strong communication skills ...

Working within a cross-functional team and reporting to a technical lead, you will operate across the machine learning development lifecycle, from data curation and synthetic data generation to model ...

Staying in sync with the latest state-of-the-art research in synthetic data generation and LLM training is key to success in this role. You will constantly lead original research initiatives through ...

next page

Showing results 1-20

Synthetic Data Generation information

See salary details

$31K

$93.2K

$169K

How much do synthetic data generation jobs pay per year?

As of Aug 17, 2026, the average yearly pay for synthetic data generation in the United States is $93,198.00, according to ZipRecruiter salary data. Most workers in this role earn between $54,500.00 and $144,500.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive in synthetic data generation?

To excel in a Synthetic Data Generation role, you need a solid background in computer science, statistics, and data science, often supported by a relevant degree and experience in machine learning. Familiarity with tools such as Python, TensorFlow, PyTorch, and synthetic data generation platforms, as well as knowledge of privacy-preserving techniques, is typically required. Strong problem-solving abilities, creativity, and effective communication set top performers apart in this field. These skills and qualities are crucial for creating high-quality, realistic synthetic datasets that support robust AI model development while safeguarding sensitive information.

What is synthetic data generation?

Synthetic data generation is the process of creating artificial datasets that mimic real-world data. This technique is used to supplement or replace actual data for purposes such as machine learning, software testing, and research, especially when real data is scarce, sensitive, or costly to obtain. Synthetic data can help improve model accuracy, protect privacy, and enable innovation by providing diverse and unbiased datasets. It is commonly used in fields like healthcare, finance, and autonomous vehicles.

What is the difference between Synthetic Data Generation vs Data Analyst?

AspectSynthetic Data GenerationData Analyst
Required CredentialsKnowledge of data science, programming, and data privacyDegree in statistics, data science, or related field
Work EnvironmentData science teams, research labs, tech companiesBusiness environments, analytics teams, consulting firms
Industry UsageAI development, machine learning, data privacyBusiness insights, reporting, decision-making
Search & Comparison IntentUnderstanding data generation techniques, privacy solutionsAnalyzing data, generating reports, insights

While Synthetic Data Generation focuses on creating artificial data for privacy and model training, Data Analysts interpret existing data to provide business insights. Both roles require data-related skills but serve different purposes within the data ecosystem.

What are the main challenges faced by professionals working in synthetic data generation, and how can they be addressed?

Professionals in synthetic data generation often encounter challenges such as ensuring the generated data accurately represents real-world scenarios while maintaining privacy and data security. Balancing realism with anonymization is crucial, especially when synthetic data is used for AI model training or testing. Collaboration with data scientists, domain experts, and privacy officers is common to validate data utility and compliance with regulations. Staying current with advances in generative models and data validation techniques also helps address these challenges and contributes to career growth in this rapidly evolving field.
More about Synthetic Data Generation jobs

What cities are hiring for Synthetic Data Generation jobs?

Cities with the most Synthetic Data Generation job openings:

What states have the most Synthetic Data Generation jobs?

States with the most job openings for Synthetic Data Generation jobs include:

Infographic showing various Synthetic Data Generation job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 84% Full Time, 11% Part Time, and 4% Contract. Highlights an 86% Physical, 4% Hybrid, and 10% Remote job distribution, with an average salary of $93,198 per year, or $44.8 per hour.

Senior Scientist, Synthetic Data Generation

NVIDIA Corporation

Santa Clara, CA • On-site

$192 - $305/hr

Other

Re-posted 6 days ago


Nvidia rating

9.6

Company rating: 9.6 out of 10

Based on 17 frontline employees who took The Breakroom Quiz

7th of 244 rated software companies


Job description

## Senior Scientist, Synthetic Data GenerationApplylocations: US, CA, Santa Clara: US, CO, Remote: US, CA, Remote: US, NY, New York: US, MA, Remotetime type: Full timeposted on: Posted Yesterdayjob requisition id: JR2019460NVIDIA is at the forefront of the AI revolution, and our research is shaping the future of large language models. We are looking for a Senior Scientist to join our team and help advance our capabilities in synthetic data generation for training frontier models. You will contribute to open-source libraries within the NVIDIA NeMo ecosystem that generate synthetic datasets across text, code, structured, and multimodal data, directly feeding the pre- and post-training of LLMs such as Nemotron. This role combines hands-on software engineering with applied research in generative methods, and you will collaborate with research, engineering, product, and model teams as well as external labs.**What you'll be doing:*** Build synthetic data generation pipelines using LLM-based methods and automated quality evaluation, producing datasets that improve the pre- and post-training of LLMs such as Nemotron — reasoning, coding, structured output, and multimodal understanding.* Advance multimodal synthetic data generation — image, document, video, and audio — in partnership with NVIDIA's model teams.* Design and maintain open-source libraries and SDKs with clean APIs and strong documentation.* Drive software excellence with modern tooling, architecture based on configuration, and professional Git/CI-CD.* Publish original research at top machine learning and AI conferences to maintain NVIDIA's technical leadership.* Mentor interns and junior researchers to develop technical growth within the team.**What we need to see:*** PhD in Computer Science, Machine Learning, Statistics, or a related field, or equivalent experience.* A research background of 3+ years in synthetic data generation, generative modeling, multimodal machine learning, or related areas. Comparable experience is also considered.* Deep technical understanding of LLMs, how data shapes their pre- and post-training, and inference frameworks such as vLLM or TGI.* Proven track record of developing or maintaining software libraries used by a broad developer community.* Strong publication record at premier venues such as NeurIPS, ICML, ICLR, ACL or similar.**Ways to stand out from the crowd:*** Open-source contributions in ML or data tooling.* Experience with multimodal generation or understanding (vision-language, document AI, video, or audio).* Building and optimizing scalable data pipelines for large-scale model training (throughput, distributed inference).* Experience generating data for agentic, tool-use, or reinforcement-learning post-training.NVIDIA is widely considered to be one of the technology world's most desirable employers. We have some of the most forward-thinking and talented people in the world working with us. If you are creative, autonomous, and passionate about building open-source tools that make AI safer and more private, we want to hear from you.Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 168,000 USD - 264,500 USD for Level 3, and 192,000 USD - 304,750 USD for Level 4.You will also be eligible for equity and benefits.Applications for this job will be accepted at least until August 10, 2026.This posting is for an existing vacancy.NVIDIA uses AI tools in its recruiting processes.NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. #J-18808-Ljbffr

What Nvidia employees say

Pay

Benefits

Hours and flexibility

Workplace

Get the full story on Breakroom


Nvidia logo

About Nvidia

Sourced by ZipRecruiter

NVIDIA has been transforming computer graphics, PC gaming, and accelerated computing for more than 25 years. It's a unique legacy of innovation that's fueled by great technology--and amazing people. Today, we're tapping into the unlimited potential of AI to define the next era of computing. An era in which our GPU acts as the brains of computers, robots, and self-driving cars that can understand the world. Doing what's never been done before takes vision, innovation, and the world's best talent.

Industry

Computer and electronic product manufacturing

Company size

10,000+ Employees

Headquarters location

Santa Clara, CA, US

Year founded

1993