1

Voice Research Task Jobs (NOW HIRING)

Speech & audio (e.g. speech enhancement, voice cloning, voice generation) * Multimodal ... LLM-as-a-Judge) together with task-specific visual, audio, and language quality metrics. Stay at ...

Improve speech recognition, voice activity detection, endpointing, and speech generation across ... accuracy, latency, naturalness, and task outcomes * Optimize end-to-end inference for ...

New

ML Research Engineer

New York, NY ยท On-site

$120K - $250K/yr

... voice. But our bigger mission goes deeper: we're building automated ontologies that model how ... Integrate AI components into autonomous agents capable of complex tasks like scheduling, order ...

... voice, video, and digital interactions in real time. Enterprises rely on Pindrop to secure billions ... acquisition/preparation tasks. * Foundational domain knowledge of biometrics, identity ...

Lead Engineer - AI Agent Voice Experience

$104K - $138K/yr

You will partner closely with applied researchers, product managers, designers, forward deployed ... Hands-on experience with ASR quality metrics such as WER and task-level evaluation methodologies.

Senior Cloud Voice Engineer

Virginia Beach, VA ยท On-site

$92K - $126K/yr

DUTIES AND TASKS: 1. Own and support enterprise voice platforms and services, including Microsoft ... Research, evaluate, and recommend modern cloud voice technologies and operating models that improve ...

... high-impact tasks and responding effectively in time-sensitive situationsKnowledge of Microsoft ... Your contributions are recognized and rewarded, and--most importantly--your voice matters. Here, yo ...

Showing results 21-40

Voice Research Task information

See salary details

$9

$26

$52

How much do voice research task jobs pay per hour?

As of Sep 6, 2026, the average hourly pay for voice research task in the United States is $26.92, according to ZipRecruiter salary data. Most workers in this role earn between $18.03 and $39.66 per hour, depending on experience, location, and employer.

What is a voice research task?

A Voice Research Task job involves recording speech samples, transcribing audio, or analyzing voice data to improve speech recognition systems, AI assistants, or linguistic models. Participants may be asked to read specific phrases, engage in conversations, or provide feedback on synthesized speech. These tasks help develop and refine voice-based technologies for better accuracy and inclusivity. No specialized skills are usually required, but clear speech and adherence to guidelines are essential.

What typical projects or tasks can I expect to work on in a voice research task?

In a Voice Research position, you may work on projects such as designing and conducting voice data collection studies, analyzing speech and audio samples, and developing or refining speech recognition systems. Responsibilities often include transcribing or annotating audio data, evaluating the performance of voice-driven technologies, and collaborating with engineers or linguists to improve product functionality. You'll likely be involved in troubleshooting research challenges and finding solutions to ensure data quality. The work environment is often collaborative, requiring ongoing coordination with cross-functional teams. This role offers exposure to the latest advancements in speech technology and opportunities for professional growth.

What are the key skills and qualifications needed to thrive in a voice research task, and why are they important?

To excel in a Voice Research role, you need a background in linguistics, phonetics, audio engineering, or a related field, with experience in collecting and analyzing voice data. Familiarity with audio editing software (like Audacity or Adobe Audition), speech processing tools, and sometimes programming languages such as Python is often required. Strong attention to detail, analytical thinking, and effective communication are valuable soft skills in this position. These abilities are crucial for producing high-quality research outputs and collaborating across multidisciplinary teams in voice technology development.

What jobs can I do using my voice?

Voice research tasks involve jobs such as voice-over artist, transcriptionist, voice actor, or virtual assistant, where clear speech and good vocal skills are essential. These roles often require recording, editing, or analyzing voice data, and may involve using specialized software or equipment. They can be performed remotely or in studio environments, with some positions requiring specific training or certifications.
More about Voice Research Task jobs

What cities are hiring for Voice Research Task jobs?

Cities with the most Voice Research Task job openings:

What are the most commonly searched types of Voice Research Task jobs?

The most popular types of Voice Research Task jobs are:

What job categories do people searching Voice Research Task jobs look for?

The top searched job categories for Voice Research Task jobs are:

Infographic showing various Voice Research Task job openings in the United States as of August 2026, with employment types broken down into 1% Internship, 1% As Needed, 84% Full Time, 11% Part Time, 1% Temporary, and 2% Contract. Highlights an 84% Physical, 4% Hybrid, and 12% Remote job distribution, with an average salary of $55,998 per year, or $26.9 per hour.

AI Research Intern

OpusClip

Mountain View, CA โ€ข On-site

Other

Re-posted 16 days ago


Job description

AI Research Intern

We're looking for an AI Research Intern to join our AI team and explore cutting-edge research across multimodal AI, LLMs, computer vision, speech, and agent systems. You'll work across OpusClip, AgentOpus, and our next-generation AI products, collaborating closely with AI researchers and engineers to investigate emerging technologies, build research prototypes, and ship features used by millions of creators worldwide.

AI Research & Model Development

  • Research and develop deep learning models in one or more of the following areas, depending on product priorities:
    • Computer vision (e.g. video enhancement, super-resolution, restoration)
    • Speech & audio (e.g. speech enhancement, voice cloning, voice generation)
    • Multimodal understanding and generation
    • LLM post-training (e.g. SFT, RLHF, DPO)

Applied AI Engineering

  • Build AI-powered product features by integrating frontier foundation models into production systems through prompt and context engineering strategies and Agent workflows (e.g., using LangChain, RAG frameworks).
  • Collaborate with product and engineering teams to rapidly prototype and ship new AI capabilities across OpusClip and AgentOpus.

Model Evaluation & Benchmarking

  • Design scalable evaluation pipelines for multimodal AI systems.
  • Develop domain-specific benchmarks using automated evaluation methods (e.g. LLM-as-a-Judge) together with task-specific visual, audio, and language quality metrics.

Stay at the Frontier

  • Keep up with the latest AI research and open-source developments.
  • Reproduce state-of-the-art research and translate new advances into production-ready systems.
Basic Qualifications
  • Education: Currently pursuing or recently completing a Master's degree in Computer Science, Artificial Intelligence, Mathematics, or a related field.
  • Deep Learning Foundation: Solid understanding of Transformer architecture and Attention mechanisms; familiarity with mainstream generative model families (GANs, diffusion models, autoregressive models).
  • Media Processing: Familiarity with media processing fundamentals (video and/or audio โ€” e.g., ffmpeg, codecs, signal processing basics).
  • Coding Skills: Strong programming skills in Python. Familiarity with Linux development environments, Git, and data structures.
  • Fluent in English with strong technical reading and writing skills, including the ability to read research papers and write technical documentation.
Hands-on Experience in One or More of the Following
  • Computer Vision (especially low-level vision): e.g., Real-ESRGAN, SwinIR, BasicVSR++, or diffusion-based SR; NTIRE / AIM challenge participation.
  • Voice / speech: voice cleaning (speech enhancement / denoising / separation), voice cloning (TTS / voice conversion), or voice generation.
  • LLM fine-tuning: SFT, RLHF / DPO, LoRA / PEFT, or post-training of open-source models.
Preferred Qualifications
  • Experience building Agent Systems or LLM-powered product features with frontier-model APIs (e.g., ChatGPT, Claude, Gemini). This role contributes to both OpusClip and AgentOpus products.
  • Familiarity with TypeScript is a bonus, helpful for shipping product features.
  • Ownership & execution: Involvement in projects from inception to completion, with strong coding fundamentals; open-source contributions are a plus.
  • Research breadth: Academic background or interest in adjacent areas โ€” video understanding and generation, multimodal systems, agents, and model evaluation / benchmarking.
  • Publications: Involvement or interest in academic research, with a focus on top-tier venues like CVPR, ICCV, ECCV, NeurIPS, ICML, ICLR, ICASSP, Interspeech, AAAI, MM, TIP, TPAMI, ACL, EMNLP etc.

Why Join OpusClip?

  • Build AI products used by millions of creators worldwide.
  • Work on cutting-edge multimodal AI, spanning LLMs, computer vision, speech, and AI agents.
  • Own projects end-to-end, from research and experimentation to production deployment.
  • Collaborate closely with experienced AI researchers and engineers in a fast-moving startup environment.
  • Opportunity to publish research while solving real-world AI problems with meaningful product impact.
  • Flexible remote/on-site internship (3 days/week required, 4+ days/week preferred).

EEO

OpusClip is proud to be an equal opportunity employer. We do not discriminate in hiring or any employment decision based on race, color, religion, national origin, age, sex (including pregnancy, childbirth, or related medical conditions), marital status, ancestry, physical or mental disability, genetic information, veteran status, gender identity or expression, sexual orientation, or other applicable legally protected characteristics. OpusClip considers qualified applicants with criminal histories, consistent with applicable federal, state and local law. Opus Clip is also committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures.