1

Speech Research Engineer Jobs (NOW HIRING)

ML Research Engineer

New York, NY ยท On-site

$120K - $250K/yr

About the Role As an ML Research Engineer at Maple, you'll be a part of our core product team ... Optimize speech recognition (ASR) , large language models (LLMs) , and text-to-speech (TTS) for ...

Senior AI Research Engineer

New York, NY ยท On-site

$114K - $157K/yr

The ideal candidate will have a proven track record as an AI research engineer, with experience across various machine learning techniques including large language models, speech models, benchmarking ...

Senior AI Research Engineer

Pittsburgh, PA ยท On-site

$101K - $139K/yr

The ideal candidate will have a proven track record as an AI research engineer, with experience across various machine learning techniques including large language models, speech models, benchmarking ...

Showing results 21-40

Speech Research Engineer information

See salary details

$37K

$106K

$142.5K

How much do speech research engineer jobs pay per year?

As of Sep 13, 2026, the average yearly pay for speech research engineer in the United States is $106,012.00, according to ZipRecruiter salary data. Most workers in this role earn between $104,000.00 and $104,000.00 per year, depending on experience, location, and employer.

What is a speech research engineer?

Speech Research Engineers are professionals who design, develop, and optimize technologies that enable computers to understand, interpret, and generate human speech. They work at the intersection of signal processing, machine learning, and linguistics to improve applications such as speech recognition, voice synthesis, and language translation systems. Their responsibilities include creating algorithms, training models on large datasets, and evaluating the performance of speech systems. These engineers often collaborate with researchers and product teams to bring cutting-edge speech technology into real-world applications.

What are the key skills and qualifications needed to thrive as a speech research engineer, and why are they important?

To thrive as a Speech Research Engineer, you need a strong background in signal processing, machine learning, and proficiency in programming languages like Python or C++, often backed by an advanced degree in computer science, electrical engineering, or a related field. Experience with frameworks such as TensorFlow or PyTorch and familiarity with speech recognition toolkits like Kaldi or HTK are typically required. Strong analytical thinking, creativity, and effective communication skills help you design innovative algorithms and collaborate in multidisciplinary teams. These skills are crucial for developing robust speech technologies and driving advancements in voice-driven applications.

What are common challenges faced by speech research engineers when working with real-world speech data?

Speech Research Engineers often encounter diverse challenges when handling real-world speech data, such as dealing with background noise, accents, dialectal variations, and inconsistent audio quality. These factors can significantly affect the accuracy and robustness of speech recognition models. Engineers must frequently preprocess and augment data, implement noise-robust algorithms, and collaborate closely with linguists and data scientists to ensure models perform well across various environments and populations.

What are popular job titles related to Speech Research Engineer jobs?

For Speech Research Engineer jobs, the most frequently searched job titles are:

Infographic showing various Speech Research Engineer job openings in the United States as of September 2026, with employment types broken down into 1% Internship, 1% As Needed, 87% Full Time, 10% Part Time, and 1% Contract. Highlights an 76% Physical, 4% Hybrid, and 20% Remote job distribution, with an average salary of $106,012 per year, or $51 per hour.

Principal Research Scientist Speech Voice Foundation Models

Alameda, CA โ€ข On-site

Other

Re-posted 15 days ago


Job description

Principal Research Scientist Speech & Audio Foundation Models

Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.


$270,000โ€“$500,000 base plus bonus, equity and benefits (US).

Relocation assistance available. Visa transfer supported

San Francisco on-site preferred | Remote considered in the US, UK and parts of Europe.

Permanent, full-time.



A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.


Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.



The role


Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.



Responsibilities



  • Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
  • Build and improve voice and speech models across TTS, STT and speech-to-speech.
  • Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.
  • Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
  • Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.



Essential



  • Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.
  • Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.
  • Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.



Desirable



  • Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
  • Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • Frontier exposure: multimodal, agents, tool use, test-time compute.
  • Public work: side projects, open-source, technical write-ups.