1

Phd Speech Jobs (NOW HIRING)

PhD level Linguist II

Burlingame, CA ยท On-site

$74K - $124K/yr

Analyze system metrics such as user opinion, lexicon transcription coverage, and Part-of-Speech ... Phd * Arabic language Job details Job ID 332928 Role PhD level Linguist II Location Burlingame ...

Showing results 21-40

Phd Speech information

See salary details

$9

$44

$67

How much do phd speech jobs pay per hour?

As of Sep 11, 2026, the average hourly pay for phd speech in the United States is $44.25, according to ZipRecruiter salary data. Most workers in this role earn between $37.74 and $50.96 per hour, depending on experience, location, and employer.

What are popular job titles related to Phd Speech jobs?

For Phd Speech jobs, the most frequently searched job titles are:

Infographic showing various Phd Speech job openings in the United States as of September 2026, with employment types broken down into 33% Full Time, and 67% Part Time. Highlights an 100% In-person job distribution, with an average salary of $92,039 per year, or $44.2 per hour.

Principal Research Scientist Speech Voice Foundation Models

Sonoma, CA โ€ข On-site

Other

Re-posted 13 days ago


Job description

Principal Research Scientist Speech & Audio Foundation Models

Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.


$270,000โ€“$500,000 base plus bonus, equity and benefits (US).

Relocation assistance available. Visa transfer supported

San Francisco on-site preferred | Remote considered in the US, UK and parts of Europe.

Permanent, full-time.



A top end research lab building realtime voice models text-to-speech, speech-to-text and speech-to-speech delivered as an API. The models run in production behind consumer applications used at very large scale, across health, learning, therapy, companionship, media and gaming.


Text-to-speech, voice cloning, speech synthesis, realtime conversational voice.



The role


Build the models that are the product! This is full-stack research ownership: you frame the question, run the experiments, and ship the result. Research is only finished when it is in production and measurable.



Responsibilities



  • Train foundation models: pre-training, reinforcement learning, reward modelling, post-training, new architectures, scaling.
  • Build and improve voice and speech models across TTS, STT and speech-to-speech.
  • Design the evaluation that proves the work: benchmarks, eval loops, LLM-as-judge, failure analysis. Evaluation is treated as a research product in its own right, not as a pre-launch checkbox.
  • Work on frontier problems adjacent to the roadmap: multimodal, agents and tool use, test-time compute.
  • Take models into production alongside the serving engineering team, inside a sub-200ms latency budget and across 100+ languages.



Essential



  • Hands-on foundation-model training. Pre-training, RL, reward modelling, post-training, scaling. Fine-tuning or building on top of someone else's model is a different discipline and is not what this role is.
  • Real voice or speech research: TTS, STT or speech-to-speech. Speech-to-speech is the strongest signal; TTS and ASR both count. Text-only research does not transfer.
  • Evidence you can point at: papers, shipped models, open-source contributions, or systems in production.



Desirable



  • Evaluation depth: benchmarks, eval loops, quality measurement, failure analysis.
  • Publications at ICML, ICLR, NeurIPS, EMNLP, ACL, AAAI, Interspeech or ICASSP.
  • PhD in ML or NLP, or equivalent practical experience you can point to.
  • Frontier exposure: multimodal, agents, tool use, test-time compute.
  • Public work: side projects, open-source, technical write-ups.