You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems -- models that don't just understand and generate high-quality speech, but hold the flow of a conversation ...
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
Director of Research, Text to Speech
San Francisco, CA · On-site +1
$213K - $328K/yr
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
Director of Research, Text to Speech
San Francisco, CA · On-site +1
$213K - $328K/yr
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
Gen AI Platform Lead
South Plainfield, NJ · On-site
$90K - $100K/yr
... to-speech systems, speech-to-text pipelines, video generation, image generation, and enterprise AI platforms. The ideal candidate will lead development of advanced AI systems including digital ...
Quick apply
Gen AI Platform Lead
South Plainfield, NJ · On-site
$90K - $100K/yr
... to-speech systems, speech-to-text pipelines, video generation, image generation, and enterprise AI platforms. The ideal candidate will lead development of advanced AI systems including digital ...
Gen AI Platform Lead
South Plainfield, NJ · On-site
$90K - $100K/yr
... to-speech systems, speech-to-text pipelines, video generation, image generation, and enterprise AI platforms. The ideal candidate will lead development of advanced AI systems including digital ...
Quick apply
Gen AI Platform Lead
South Plainfield, NJ · On-site
$90K - $100K/yr
... to-speech systems, speech-to-text pipelines, video generation, image generation, and enterprise AI platforms. The ideal candidate will lead development of advanced AI systems including digital ...
Gen AI Platform Lead
South Plainfield, NJ · On-site
$90K - $100K/yr
... to-speech systems, speech-to-text pipelines, video generation, image generation, and enterprise AI platforms. The ideal candidate will lead development of advanced AI systems including digital ...
Quick apply
Gen AI Platform Lead
South Plainfield, NJ · On-site
$90K - $100K/yr
... to-speech systems, speech-to-text pipelines, video generation, image generation, and enterprise AI platforms. The ideal candidate will lead development of advanced AI systems including digital ...
Expertise in speech-to-text, text-to-speech, and designing natural, human-like voice conversations. • NLP & Machine Learning: Strong background in Natural Language Processing, intent recognition ...
Expertise in speech-to-text, text-to-speech, and designing natural, human-like voice conversations. • NLP & Machine Learning: Strong background in Natural Language Processing, intent recognition ...
Whether text-to-text, text-to-speech, speech-to-text, or speech-to-speech, machine translation can help overlooked communities finally be understood in the world. HLT will bring critical educational ...
Whether text-to-text, text-to-speech, speech-to-text, or speech-to-speech, machine translation can help overlooked communities finally be understood in the world. HLT will bring critical educational ...
Information Technology_USA - USA_Engineer
Jacksonville, FL · On-site
$49 - $67/hr
... text/text-to-speech. Role Descriptions: 4+ years of commercial software development experience.Design and implement scalable CCaaS and IVA solutions leveraging leading Cloud and enterprise ...
Information Technology_USA - USA_Engineer
Jacksonville, FL · On-site
$49 - $67/hr
... text/text-to-speech. Role Descriptions: 4+ years of commercial software development experience.Design and implement scalable CCaaS and IVA solutions leveraging leading Cloud and enterprise ...
Transcription Specialist
Phoenix, AZ · On-site
You'll listen to English-language audio and correct the corresponding speech-to-text transcript according to the project SOP. You'll also help track and report Word Error Rate (WER) as part of the ...
Quick apply
Transcription Specialist
Phoenix, AZ · On-site
You'll listen to English-language audio and correct the corresponding speech-to-text transcript according to the project SOP. You'll also help track and report Word Error Rate (WER) as part of the ...
The Opportunity We're hiring for a Product Manager focused on AI speech (text-to-speech, speech-to-text, voice cloning, real-time voice, and more). You'll work closely with us to develop and enhance ...
The Opportunity We're hiring for a Product Manager focused on AI speech (text-to-speech, speech-to-text, voice cloning, real-time voice, and more). You'll work closely with us to develop and enhance ...
Ex-JPMC AI/ML Engineer
Plano, TX · On-site
... speech-to-text, text-to-speech) · Design and maintain APIs (REST, GraphQL), microservices, and event-driven architectures · Own CI/CD pipelines, containerization (Docker, ECS/EKS), and ...
Ex-JPMC AI/ML Engineer
Plano, TX · On-site
... speech-to-text, text-to-speech) · Design and maintain APIs (REST, GraphQL), microservices, and event-driven architectures · Own CI/CD pipelines, containerization (Docker, ECS/EKS), and ...
Deepgram is the leading platform in the Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech solutions. They are seeking a Software Engineer to join their new business ...
Deepgram is the leading platform in the Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech solutions. They are seeking a Software Engineer to join their new business ...
Deepgram is the leading voice AI platform for developers building speech-to-text and text-to-speech offerings. They are seeking a Software Engineer to join their new business unit focused on ...
Deepgram is the leading voice AI platform for developers building speech-to-text and text-to-speech offerings. They are seeking a Software Engineer to join their new business unit focused on ...
Senior Software Engineer - Model Evaluation & AI Systems
$180K - $240K/yr
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
Senior Software Engineer - Model Evaluation & AI Systems
$180K - $240K/yr
Company Overview Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT), text-to-speech (TTS), and building ...
Speech To Text Proofreader information
See salary details
$13.70 - $16.46
9% of jobs
$16.46 - $19.21
13% of jobs
$20.05 is the 25th percentile. Wages below this are outliers.
$19.21 - $21.96
9% of jobs
$21.96 - $24.72
6% of jobs
The median wage is $26.98 / hr.
$24.72 - $27.47
15% of jobs
$27.47 - $30.22
17% of jobs
$32.03 is the 75th percentile. Wages above this are outliers.
$30.22 - $32.98
8% of jobs
$32.98 - $35.73
8% of jobs
$35.73 - $38.48
7% of jobs
$38.48 - $41.24
2% of jobs
$41.24 - $43.99
4% of jobs
$13
$27
$43
How much do speech to text proofreader jobs pay per hour?
What are popular job titles related to Speech To Text Proofreader jobs?
For Speech To Text Proofreader jobs, the most frequently searched job titles are:

Artificial Intelligence Researcher
Sunnyvale, CA • On-site
Other
This job post has expired 1 day ago. Applications are no longer accepted.
Job description
Kotoba's speech models are licensed to Fortune 50 companies and US big tech, and power an app reaching 2,000–3,000 new users a day. We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI.
Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems — models that don't just understand and generate high-quality speech, but hold the flow of a conversation: turn-taking, interruptions, overlapping speech, backchannels, response timing, prosody, and latency. Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment.
■ About Kotoba
Kotoba is a generative AI company on a mission to become the default for voice AI in East Asia. At our core is a low-latency, high-accuracy speech translation model that connects conversations so naturally it feels as though both speakers share the same language, supporting Japanese, English, Korean, Chinese, Spanish, and other major language pairs. We also build ultra-low-latency speech-to-text and text-to-speech models that run everywhere from the data center to edge devices, and we license this foundational technology to Fortune 50 companies and major US tech firms. We work from two hubs: Tokyo and San Francisco.
Our own product, the Kotoba app, is available on iOS and Android. Since launch it has grown to a steady 2,000–3,000 new downloads per day and reached No. 1 in its App Store and Google Play category, ahead of the likes of Google Translate. Enterprise adoption is accelerating in Japan, and the app has supported nearly 100 live events including SusHi Tech Tokyo.
Kotoba was founded in 2023 by two Japanese generative AI researchers with PhDs from top US universities. We've raised over ¥3 billion (roughly US$23M) from prominent VCs in Japan and the US — including Kindred Ventures and Globis Capital Partners — and from the corporate venture arms of leading US and Japanese enterprises. We also receive strong government support in Japan for AI model training.
■ What you'll do
- Define and execute research projects for next-generation voice AI across speech-to-speech, speech-to-text, and text-to-speech systems
- Develop full-duplex conversational models that listen and speak simultaneously while handling turn-taking, interruptions, overlapping speech, backchannels, and end-of-turn prediction
- Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models
- Conduct multilingual and cross-lingual research, particularly for Japanese, Korean, Chinese, English, and other languages central to our products
- Explore architectures that orchestrate speech, language, reasoning, retrieval, and tool-use models behind a unified real-time voice interface
- Build and scale model training and inference pipelines on distributed GPU infrastructure, optimizing models for low-latency deployment
- Work with research, product, and infrastructure engineers to move promising research into our applications, APIs, SDKs, and customer projects
■ What we're looking for
Required
- A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field
- A strong research track record, demonstrated through publications at leading conferences or journals in machine learning, speech, NLP, or related areas
- Deep expertise in at least one relevant area: speech-to-speech modeling, speech translation, spoken dialogue systems, speech recognition, speech generation, multimodal foundation models, large language models, or AI model orchestration
- Hands-on experience designing, implementing, training, and evaluating modern neural models in PyTorch or JAX
- Strong knowledge of modern speech and language architectures, including transformers, streaming models, autoregressive and non-autoregressive models, and foundation-model training
- The ability to formulate original research questions, design rigorous experiments, analyze results critically, and turn promising ideas into working systems
- Familiarity with large-scale model training, inference, data pipelines, distributed computing, and GPU-based experimentation
- Strong written and verbal communication, including professional proficiency in English
Preferred
- Research experience in full-duplex speech-to-speech, speech recognition, or speech generation — particularly turn-taking, interruptions, backchannels, dialogue timing, or conversational fluency
- Research experience involving Japanese, Korean, Chinese, or other East Asian languages, including multilingual or cross-lingual modeling
- Knowledge of audio tokenization, neural audio codecs, streaming speech recognition, streaming speech generation, or low-latency speech architectures
- Experience with distributed training and efficient inference for large speech, language, or multimodal models
- Research experience with systems that orchestrate multiple models, agents, retrieval components, reasoning modules, or external tools
- Previous experience at an industrial research lab, major AI organization, technology company, or research-driven startup, particularly transferring research into production
- A record of open-source contributions
■ Location
San Francisco
About Kotoba
Sourced by ZipRecruiter
Industry
Translation services
Company size
1 - 10 Employees
Headquarters location
Bethesda, MD, US