1

Synthetic Data Generation Jobs in Florida (NOW HIRING)

Develop synthetic data generation pipelines for training perception, manipulation, and navigation models. We're Looking For: * BS, MS, or PhD in Robotics, Computer Science, Mechanical Engineering, or ...

Direct experience generating synthetic data for model training, including sensor simulation, annotation pipelines, and large-scale dataset generation. * Working knowledge of OpenUSD as a robotics ...

Principal Software Engineer (Python)

Sunrise, FL · On-site +1

$128K - $172K/yr

Monitor and provide recommendations on the rapidly evolving AI landscape, including SLM success areas, prompt engineering, and synthetic data generation. * Mentor Senior & Junior Engineers: Act as a ...

Senior ML Engineer

Dania Beach, FL

$102K - $141K/yr

Knowledge of synthetic data generation techniques for augmenting fine-tuning datasets. * Exposure to multimodal models (vision-language, speech-language) or voice/speech AI systems. * Contributions ...

Machine Learning Engineer

Orlando, FL · On-site

$120 - $160/hr

Kubernetes, Docker, APIGEE, Terraform Data: Vector databases (Pinecone, Weaviate), feature stores ... Build text‑to‑image and text‑to‑video generation systems * Develop speech synthesis and ...

New

... generation, and engagement features, helping you save prep time and focus on impactful teaching ... Guides students through designing multistep synthesis routes, predicting regiochemistry and ...

next page

Showing results 1-20

Synthetic Data Generation information

What are the key skills and qualifications needed to thrive in synthetic data generation?

To excel in a Synthetic Data Generation role, you need a solid background in computer science, statistics, and data science, often supported by a relevant degree and experience in machine learning. Familiarity with tools such as Python, TensorFlow, PyTorch, and synthetic data generation platforms, as well as knowledge of privacy-preserving techniques, is typically required. Strong problem-solving abilities, creativity, and effective communication set top performers apart in this field. These skills and qualities are crucial for creating high-quality, realistic synthetic datasets that support robust AI model development while safeguarding sensitive information.

What is synthetic data generation?

Synthetic data generation is the process of creating artificial datasets that mimic real-world data. This technique is used to supplement or replace actual data for purposes such as machine learning, software testing, and research, especially when real data is scarce, sensitive, or costly to obtain. Synthetic data can help improve model accuracy, protect privacy, and enable innovation by providing diverse and unbiased datasets. It is commonly used in fields like healthcare, finance, and autonomous vehicles.

What is the difference between Synthetic Data Generation vs Data Analyst?

AspectSynthetic Data GenerationData Analyst
Required CredentialsKnowledge of data science, programming, and data privacyDegree in statistics, data science, or related field
Work EnvironmentData science teams, research labs, tech companiesBusiness environments, analytics teams, consulting firms
Industry UsageAI development, machine learning, data privacyBusiness insights, reporting, decision-making
Search & Comparison IntentUnderstanding data generation techniques, privacy solutionsAnalyzing data, generating reports, insights

While Synthetic Data Generation focuses on creating artificial data for privacy and model training, Data Analysts interpret existing data to provide business insights. Both roles require data-related skills but serve different purposes within the data ecosystem.

What are the main challenges faced by professionals working in synthetic data generation, and how can they be addressed?

Professionals in synthetic data generation often encounter challenges such as ensuring the generated data accurately represents real-world scenarios while maintaining privacy and data security. Balancing realism with anonymization is crucial, especially when synthetic data is used for AI model training or testing. Collaboration with data scientists, domain experts, and privacy officers is common to validate data utility and compliance with regulations. Staying current with advances in generative models and data validation techniques also helps address these challenges and contributes to career growth in this rapidly evolving field.
What are popular job titles related to Synthetic Data Generation jobs in Florida? For Synthetic Data Generation jobs in Florida, the most frequently searched job titles are:
What job categories do people searching Synthetic Data Generation jobs in Florida look for? The top searched job categories for Synthetic Data Generation jobs in Florida are:
What cities in Florida are hiring for Synthetic Data Generation jobs? Cities in Florida with the most Synthetic Data Generation job openings:
Infographic showing various Synthetic Data Generation job openings in Florida as of July 2026, with employment types broken down into 1% As Needed, 81% Full Time, 9% Part Time, 1% Temporary, and 8% Contract. Highlights an 81% Physical, 3% Hybrid, and 16% Remote job distribution.

Robotics Data Pipeline Intern

Persona AI

Pensacola, FL • On-site

Full-time

Re-posted 24 days ago


Job description

Robotics Data Pipeline Intern - Multimodal Data
About Us At Persona, we're building the next generation of humanoid robots, and that requires an unprecedented volume of high-quality, multimodal data. We're moving beyond basic teleoperation to leverage massive datasets of in-the-wild egocentric video combined with dense sensor streams (IMU, haptics, kinematics, and high-fidelity force profiles). We're looking for a curious, technically sharp intern to roll up their sleeves and help us turn raw, unstructured multimodal data into high-fidelity training assets for our robots.
The Role As a Data Pipeline Intern, you'll work directly alongside our data and robotics engineering teams to support the infrastructure that feeds our foundation models. You'll get hands-on experience with real multimodal data challenges, from sensor stream processing and video pipeline optimization to force analysis and kinematic retargeting. This is not a "fetch coffee and shadow engineers" internship. You'll own real work and ship real code.
What You'll Work On
  • Rebuilding and extending pipelines that ingest and synchronously process egocentric video alongside rich sensor streams (IMU, force-torque, tactile, proprioception)
  • Owning post-processing algorithms for force analysis and hidden state inference, including contact force estimation, occlusion handling, and inverse kinematics gap-filling
  • Bridging kinematic retargeting work that translates human hand tracking into humanoid end-effector coordinates
  • Optimizing and testing data augmentation strategies (spatial, temporal, synthetic viewpoints, sensor noise injection)
  • Tying together work across our Hardware Teleoperation Team to help align human-robot play-data across modalities

What We're Looking For
  • Currently pursuing a B.S., M.S., or Ph.D. in Computer Science, Data Engineering, Machine Learning, Robotics, or a related field
  • Solid Python skills and exposure to PyTorch, particularly around data loading or multimodal datasets
  • Coursework or project experience with computer vision, time-series data, or sensor processing
  • Familiarity with video processing tools (OpenCV, FFmpeg) or pose estimation frameworks (MediaPipe) is a plus
  • Awareness of imitation learning, VLA architectures, or human-to-robot transfer concepts is a plus, but genuine curiosity counts for a lot here

Bonus Points
  • Experience with NVIDIA's robotics stack (Isaac, Cosmos, GR00T)
  • Exposure to distributed computing (Ray, Spark) or simulation environments (Omniverse, MuJoCo)
  • Any project work involving synthetic data generation or tactile/spatial data representations