1

Data Generator User Jobs in New York (NOW HIRING)

Founding AI Engineer

Manhattan, NY · On-site

$200 - $250/hr

Layered agents, generator and critic loops, planners and executors, and orchestration where one ... data and can act on it. Building AI products that ask very little of the user. This means the UI ...

Data Generator User information

What is a Data Generator User?

A Data Generator User is someone who utilizes data generation tools or software to create synthetic or simulated data sets. These users often work in data science, software testing, or machine learning to generate data for testing algorithms, validating systems, or training models when real-world data is unavailable or sensitive. Data Generator Users are skilled at configuring parameters and ensuring the generated data meets specific requirements, such as volume, variety, and format. Their work helps organizations test and develop robust data-driven solutions without compromising privacy or relying solely on existing data.

What are the key skills and qualifications needed to thrive as a Data Generator User, and why are they important?

To thrive as a Data Generator User, you need a strong understanding of data analysis, attention to detail, and familiarity with data management principles, typically supported by relevant education or experience in data handling. Proficiency in data generation tools, spreadsheet software, and sometimes programming languages like Python or SQL is often required. Strong problem-solving abilities, communication skills, and the ability to work independently are key soft skills in this role. These skills ensure accurate, efficient, and reliable data creation, which is critical for supporting business analytics and decision-making processes.

What are some typical challenges faced by Data Generator Users when ensuring data quality and consistency?

One common challenge Data Generator Users encounter is maintaining high data quality and consistency, especially when generating large datasets for testing or analytics. Ensuring that generated data accurately reflects real-world scenarios and edge cases requires careful planning and validation. Additionally, collaborating with development, QA, and analytics teams to understand their specific data requirements can be complex but is essential for delivering valuable datasets. Regular reviews and automated validation checks can help minimize errors and ensure reliable results.

What is the difference between Data Generator User vs Data Analyst?

AspectData Generator UserData Analyst
Required CredentialsBasic technical skills, familiarity with data toolsDegree in statistics, data science, or related field
Work EnvironmentUses data generation tools in various industriesAnalyzes data to derive insights, often in office settings
Employer & Industry UsageTech companies, research labs, data firmsBusiness, finance, healthcare, marketing

The main difference is that Data Generator Users focus on creating synthetic data using specialized tools, while Data Analysts interpret and analyze existing data to support decision-making. Both roles require technical skills, but Data Analysts typically have more advanced certifications and focus on data interpretation.

What are popular job titles related to Data Generator User jobs in New York?

For Data Generator User jobs in New York, the most frequently searched job titles are:

What job categories do people searching Data Generator User jobs in New York look for?

The top searched job categories for Data Generator User jobs in New York are:

What cities in New York are hiring for Data Generator User jobs?

Cities in New York with the most Data Generator User job openings:

Senior Software Engineer, AI/ML, NotebookLM Content Studio

Socket.dev

Manhattan, NY • On-site

$174 - $253/hr

Other

This job post has expired 1 day ago. Applications are no longer accepted.


Job description

MINIMUM QUALIFICATIONS:

  • Bachelor’s degree in Computer Science, Mathematics, or equivalent practical experience.

  • 5 years of experience with software development in programming languages like Python, C++, or Java.

  • 3 years of experience testing, maintaining, or launching software products, and 1 year of experience with software design and architecture.

  • 3 years of experience with ML infrastructure (including model deployment, evaluation, optimization, and debugging).

  • 3 years of experience with speech/audio (e.g., voice synthesis), Generative AI/multimodal media synthesis, reinforcement learning, or another ML field.


PREFERRED QUALIFICATIONS:

  • Master\'s degree or PhD in Computer Science, Machine Learning, or related technical field.

  • 5 years of experience with advanced data structures, algorithms, and machine learning optimization.

  • 1 year of experience in a technical leadership role, guiding rapid prototyping or directing product-driven engineering efforts.

  • Experience developing highly accessible consumer-facing technologies.

  • Experience with advanced prompt engineering, creative ML applications, and building consumer products that utilize Large Language Models (LLMs) or multimodal media formats (such as video, audio, or image synthesis).


ABOUT THE JOB:

Google\'s software engineers develop the next-generation technologies that change how billions of users connect, explore, and interact with information and one another. Our products need to handle information at massive scale, and extend well beyond web search. We\'re looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. As a software engineer, you will work on a specific project critical to Google’s needs with opportunities to switch teams and projects as you and our fast-paced business grow and evolve. We need our engineers to be versatile, display leadership qualities and be enthusiastic to take on new problems across the full-stack as we continue to push technology forward.


Within Labs, the Language Applications team manages a high-impact portfolio, including AI Studio, the Gemini API, and NotebookLM. The role is within GemFM, the team responsible for multimodal artifact generations.


Our focus is on synthetic audio, video, code, and any multimodal generation. We are the team that built audio overviews and cinematic video overviews, and we are seeking passionate ML SWEs to help us expand the boundaries of creative content generation, redefining how and where these new forms of media are made.


Labs is a group focused on incubating early-stage efforts in support of Google’s mission to organize the world’s information and make it universally accessible and useful. Our team exists to help discover and create new ways to advance our core products through exploration and the application of new technologies. We work to build new solutions that have the potential to transform how users interact with Google. Our goal is to drive innovation by developing new Google products and capabilities that deliver significant impact over longer timeframes.


Individual pay is determined by factors including job-related skills, experience, and relevant education or training.


US: $174000 - $253000 (USD) + 15% bonus target + equity + benefits


Learn more about benefits at Google


[https://www.google.com/about/careers/applications/benefits/].


RESPONSIBILITIES:

  • Write and test product or system development code. Write, test, and maintain production-ready code for NotebookLM’s Content Studio, focusing on multimodal media pipelines (synthetic audio, video, and code).

  • Collaborate with peers and stakeholders through design and code reviews. Partner with engineering, UX, and research teams to ensure best practices in prompt engineering, model safety, and system efficiency.

  • Contribute to documentation and educational content. Create and update technical documentation and guidelines, adapting content to keep pace with model upgrades and user feedback.

  • Diagnose and resolve pipeline anomalies within the GemFM generator, optimizing for TPU/GPU performance, network latency, and output quality.

  • Design and deploy advanced solutions in specialized ML fields (such as voice synthesis or long-context LLMs) using internal ML infrastructure.

#J-18808-Ljbffr