Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
Take models into production alongside the serving engineering team, inside a sub-200ms latency ... Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest ...
... speech (TTS). They are seeking a highly skilled Machine Learning Engineer to join their Research ... team, where the role involves prototyping and validating novel modeling ideas and scaling them ...
... speech (TTS). They are seeking a highly skilled Machine Learning Engineer to join their Research ... team, where the role involves prototyping and validating novel modeling ideas and scaling them ...
About the Role We are hiring a Research Engineer to accelerate Smule's research teams by building ... Experience with audio, speech, or music signal processing pipelines. * Familiarity with ...
Quick apply
About the Role We are hiring a Research Engineer to accelerate Smule's research teams by building ... Experience with audio, speech, or music signal processing pipelines. * Familiarity with ...
About the Role We are hiring a Research Engineer to accelerate Smule's research teams by building ... Experience with audio, speech, or music signal processing pipelines. * Familiarity with ...
About the Role We are hiring a Research Engineer to accelerate Smule's research teams by building ... Experience with audio, speech, or music signal processing pipelines. * Familiarity with ...
About the Role We are hiring a Research Engineer to accelerate Smule's research teams by building ... Experience with audio, speech, or music signal processing pipelines. * Familiarity with ...
About the Role We are hiring a Research Engineer to accelerate Smule's research teams by building ... Experience with audio, speech, or music signal processing pipelines. * Familiarity with ...
Principal Research & Engineering, Realtime Voice AI
$400K - $550K/yr
Determine build-vs-buy-vs-train strategies for core audio, speech, and realtime interaction components. * Direct research and engineering efforts focused on speech quality, naturalness ...
Principal Research & Engineering, Realtime Voice AI
$400K - $550K/yr
Determine build-vs-buy-vs-train strategies for core audio, speech, and realtime interaction components. * Direct research and engineering efforts focused on speech quality, naturalness ...
AI Research Engineer- Speech 1
Redmond, WA ยท On-site
$160K/yr
D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning. * 2+ years of industry or applied research experience in ...
AI Research Engineer- Speech 1
Redmond, WA ยท On-site
$160K/yr
D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning. * 2+ years of industry or applied research experience in ...
ML Research Engineer
New York, NY ยท On-site
$120K - $250K/yr
About the Role As an ML Research Engineer at Maple, you'll be a part of our core product team ... Optimize speech recognition (ASR) , large language models (LLMs) , and text-to-speech (TTS) for ...
ML Research Engineer
New York, NY ยท On-site
$120K - $250K/yr
About the Role As an ML Research Engineer at Maple, you'll be a part of our core product team ... Optimize speech recognition (ASR) , large language models (LLMs) , and text-to-speech (TTS) for ...
... research in speech technology and machine learning. As a member of the team, you will be inspired by a diversity of challenging problems, collaborate with world-class machine learning engineers and ...
... research in speech technology and machine learning. As a member of the team, you will be inspired by a diversity of challenging problems, collaborate with world-class machine learning engineers and ...
Senior AI Research Engineer
New York, NY ยท On-site
$114K - $157K/yr
The ideal candidate will have a proven track record as an AI research engineer, with experience across various machine learning techniques including large language models, speech models, benchmarking ...
Senior AI Research Engineer
New York, NY ยท On-site
$114K - $157K/yr
The ideal candidate will have a proven track record as an AI research engineer, with experience across various machine learning techniques including large language models, speech models, benchmarking ...
Senior AI Research Engineer
Pittsburgh, PA ยท On-site
$101K - $139K/yr
The ideal candidate will have a proven track record as an AI research engineer, with experience across various machine learning techniques including large language models, speech models, benchmarking ...
Senior AI Research Engineer
Pittsburgh, PA ยท On-site
$101K - $139K/yr
The ideal candidate will have a proven track record as an AI research engineer, with experience across various machine learning techniques including large language models, speech models, benchmarking ...
Senior Machine Learning Research Engineer
San Francisco, CA ยท On-site
$123K - $169K/yr
Senior Machine Learning Research Engineer San Francisco, United States | Posted on 09/03/2026 AI ... Publishing and advancing the state of the art in speech, audio, and multimodal ML (NeurIPS, ICML ...
Senior Machine Learning Research Engineer
San Francisco, CA ยท On-site
$123K - $169K/yr
Senior Machine Learning Research Engineer San Francisco, United States | Posted on 09/03/2026 AI ... Publishing and advancing the state of the art in speech, audio, and multimodal ML (NeurIPS, ICML ...
Research Engineer - 530162
Tuscaloosa, AL ยท On-site
$56K - $73K/yr
Research Engineer - 530162 Job no: 530162 Work type: Regular Full-time (Benefits eligible) Location ... neutrality, free speech, and academic freedom. Advertised: 04 Sep 2026 Central Daylight Time ...
Research Engineer - 530162
Tuscaloosa, AL ยท On-site
$56K - $73K/yr
Research Engineer - 530162 Job no: 530162 Work type: Regular Full-time (Benefits eligible) Location ... neutrality, free speech, and academic freedom. Advertised: 04 Sep 2026 Central Daylight Time ...
Real-Time Speech & Audio Research Engineer
San Diego, CA ยท On-site
$148K - $222K/yr
Qualcomm is looking for a candidate for its Speech team to work on transforming communication ... The role requires executing fundamental research and developing technologies for real-time ...
Real-Time Speech & Audio Research Engineer
San Diego, CA ยท On-site
$148K - $222K/yr
Qualcomm is looking for a candidate for its Speech team to work on transforming communication ... The role requires executing fundamental research and developing technologies for real-time ...
Speech Research Engineer information
See salary details
$37K - $46.6K
3% of jobs
$46.6K - $56.2K
0% of jobs
$56.2K - $65.8K
0% of jobs
$65.8K - $75.4K
1% of jobs
$75.4K - $85K
4% of jobs
$85K - $94.5K
4% of jobs
$98.5K is the 25th percentile. Wages below this are outliers.
$94.5K - $104.1K
30% of jobs
The median wage is $105.8K / yr.
$104.1K - $113.7K
43% of jobs
$113.7K - $123.3K
6% of jobs
$123.3K - $132.9K
1% of jobs
$132.9K - $142.5K
7% of jobs
$37K
$106K
$142.5K
How much do speech research engineer jobs pay per year?
What is a speech research engineer?
What are the key skills and qualifications needed to thrive as a speech research engineer, and why are they important?
What are common challenges faced by speech research engineers when working with real-world speech data?
What are popular job titles related to Speech Research Engineer jobs?
For Speech Research Engineer jobs, the most frequently searched job titles are:

Principal Research Scientist Speech Voice Foundation Models
Alameda, CA โข On-site
Other
Re-posted 15 days ago
Job description
Principal Research Scientist Speech & Audio Foundation Models
Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.
$270,000โ$500,000 base plus bonus, equity and benefits (US).
Relocation assistance available. Visa transfer supported
San Francisco on-site preferred | Remote considered in the US, UK and parts of Europe.
Permanent, full-time.
A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.
Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.
The role
Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.
Responsibilities
- Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
- Build and improve voice and speech models across TTS, STT and speech-to-speech.
- Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.
- Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
- Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.
Essential
- Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.
- Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.
- Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.
Desirable
- Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
- Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
- PhD in ML or NLP, or equivalent practical experience you can point to.
- Frontier exposure: multimodal, agents, tool use, test-time compute.
- Public work: side projects, open-source, technical write-ups.