Hi, We are specifically seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS ... Collaborate with cross-functional teams to develop and deliver new speech AI features.
Hi, We are specifically seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS ... Collaborate with cross-functional teams to develop and deliver new speech AI features.
... AI technologies. This role focuses on training and optimizing speech models, improving model ... The ideal candidate has strong experience in speech synthesis (Text-to-Speech) or Speech-to-Text ...
... AI technologies. This role focuses on training and optimizing speech models, improving model ... The ideal candidate has strong experience in speech synthesis (Text-to-Speech) or Speech-to-Text ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) systems capable of producing ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) systems capable of producing ...
Lead AI/ML Engineer
$170K - $190K/yr
You will lead the design and delivery of end-to-end voice AI solutions, combining large language models with speech technologies such as speech-to-text, text-to-speech, and real-time streaming audio ...
Quick apply
Lead AI/ML Engineer
$170K - $190K/yr
You will lead the design and delivery of end-to-end voice AI solutions, combining large language models with speech technologies such as speech-to-text, text-to-speech, and real-time streaming audio ...
Maintain and enhance text to speech evaluation systems * Analyze model accuracy and bias and recommend improvements * Improve processes related to speech data preparation, augmentation, and filtering
Quick apply
Maintain and enhance text to speech evaluation systems * Analyze model accuracy and bias and recommend improvements * Improve processes related to speech data preparation, augmentation, and filtering
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT) and text-to-speech (TTS). They are seeking a highly ...
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT) and text-to-speech (TTS). They are seeking a highly ...
ML Researcher, Speech
San Francisco, CA · On-site
$200K - $250K/yr
... to-speech conversational AI model that understands and responds like a human, in real time. You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech ...
ML Researcher, Speech
San Francisco, CA · On-site
$200K - $250K/yr
... to-speech conversational AI model that understands and responds like a human, in real time. You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech ...
Audio Solutions Architect
$150K - $230K/yr
Innodata builds the high-quality voice and audio datasets that power the world's leading speech AI - text-to-speech, speech recognition, and the new generation of speech-to-speech and conversational ...
Audio Solutions Architect
$150K - $230K/yr
Innodata builds the high-quality voice and audio datasets that power the world's leading speech AI - text-to-speech, speech recognition, and the new generation of speech-to-speech and conversational ...
Lead Engineer - AI Agent Voice Experience
$104K - $138K/yr
Cresta's unified AI platform combines conversational AI agents, real-time human agent augmentation ... text to speech, speech-to-speech models, and production optimization. You will partner closely with ...
Lead Engineer - AI Agent Voice Experience
$104K - $138K/yr
Cresta's unified AI platform combines conversational AI agents, real-time human agent augmentation ... text to speech, speech-to-speech models, and production optimization. You will partner closely with ...
Lead Engineer - AI Agent Voice Experience
OR · On-site +1
... text to speech, speech-to-speech models, and production optimization. You will partner closely with ... Improve the quality of voice AI systems through error analysis, data curation, metric design ...
Lead Engineer - AI Agent Voice Experience
OR · On-site +1
... text to speech, speech-to-speech models, and production optimization. You will partner closely with ... Improve the quality of voice AI systems through error analysis, data curation, metric design ...
Liquid AI is a company spun out of MIT CSAIL that builds general-purpose AI systems. They are ... speech and text-to-text training, including synthetic dialogue, function calling examples, and ...
Liquid AI is a company spun out of MIT CSAIL that builds general-purpose AI systems. They are ... speech and text-to-text training, including synthetic dialogue, function calling examples, and ...
Member of Technical Staff - Audio and Voice AI
New York, NY · On-site
$220K - $320K/yr
Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial ...
Member of Technical Staff - Audio and Voice AI
New York, NY · On-site
$220K - $320K/yr
Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial ...
Senior Machine Learning Engineer, Voice AI
San Francisco, CA · On-site
$144K - $190K/yr
Our Voice AI platform powers production-grade, real-time voice agents and applications - serving speech-to-text and text-to-speech models with best-in-class latency and reliability. We're looking for ...
Senior Machine Learning Engineer, Voice AI
San Francisco, CA · On-site
$144K - $190K/yr
Our Voice AI platform powers production-grade, real-time voice agents and applications - serving speech-to-text and text-to-speech models with best-in-class latency and reliability. We're looking for ...
Member of Technical Staff -- Audio and Voice AI
New York, NY · On-site
$220K - $320K/yr
Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial ...
Quick apply
Member of Technical Staff -- Audio and Voice AI
New York, NY · On-site
$220K - $320K/yr
Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial ...
Senior Machine Learning Engineer, Voice AI
San Francisco, CA · On-site
$144K - $190K/yr
Our Voice AI platform powers production-grade, real-time voice agents and applications - serving speech-to-text and text-to-speech models with best-in-class latency and reliability. We're looking for ...
Senior Machine Learning Engineer, Voice AI
San Francisco, CA · On-site
$144K - $190K/yr
Our Voice AI platform powers production-grade, real-time voice agents and applications - serving speech-to-text and text-to-speech models with best-in-class latency and reliability. We're looking for ...
Member of Technical Staff -- Audio and Voice AI
San Francisco, CA · On-site
$220K - $320K/yr
Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial ...
Quick apply
Member of Technical Staff -- Audio and Voice AI
San Francisco, CA · On-site
$220K - $320K/yr
Craft High-Quality Voice UX: use modern speech-to-text, text-to-speech, and conversational AI platforms to create natural, responsive, and emotionally aware voice experiences tailored to financial ...
Software Engineer - Applied AI (Senior or Staff Level)
Manhattan, NY · On-site
$134K - $177K/yr
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech. They are seeking a Software Engineer to ...
Software Engineer - Applied AI (Senior or Staff Level)
Manhattan, NY · On-site
$134K - $177K/yr
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text and text-to-speech. They are seeking a Software Engineer to ...
Voice AI
Alpharetta, GA · On-site
... Text-to-Speech (TTS). • Setting up workflow automations for lead qualification, booking ... Required : • Familiarity with AI workflows • Natural language processing (NLP) pipelines • ...
Voice AI
Alpharetta, GA · On-site
... Text-to-Speech (TTS). • Setting up workflow automations for lead qualification, booking ... Required : • Familiarity with AI workflows • Natural language processing (NLP) pipelines • ...
Research Scientist-Voice and Audio Ai
$225K - $400K/yr
... AI. As a Founding Machine Learning Research Engineer at Retell, you'll focus on advancing model ... LLM pre-training and post training, transcription model training, text to speech model training, or ...
Quick apply
Research Scientist-Voice and Audio Ai
$225K - $400K/yr
... AI. As a Founding Machine Learning Research Engineer at Retell, you'll focus on advancing model ... LLM pre-training and post training, transcription model training, text to speech model training, or ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) technology capable of producing ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) technology capable of producing ...
Ai Text To Speech information
See salary details
$15.63 - $20.54
0% of jobs
$20.54 - $25.46
2% of jobs
$25.46 - $30.38
6% of jobs
$30.38 - $35.29
12% of jobs
$36.60 is the 25th percentile. Wages below this are outliers.
$35.29 - $40.21
17% of jobs
The median wage is $43.54 / hr.
$40.21 - $45.13
18% of jobs
$45.13 - $50.04
17% of jobs
$50.76 is the 75th percentile. Wages above this are outliers.
$50.04 - $54.96
13% of jobs
$54.96 - $59.88
8% of jobs
$59.88 - $64.79
4% of jobs
$64.79 - $69.71
2% of jobs
$15
$43
$69
How much do ai text to speech jobs pay per hour?
What is the difference between Ai Text To Speech vs Voice Actor?
| Aspect | Ai Text To Speech | Voice Actor |
|---|---|---|
| Credentials | None required, but technical skills helpful | Voice training, acting skills, often a demo reel |
| Work Environment | Remote, software-based | Studio or on-location recording |
| Industry Usage | Technology, media, customer service | Entertainment, advertising, narration |
| Work Nature | Automated voice generation, programming | Live or pre-recorded voice performances |
Ai Text To Speech involves using software to convert text into synthetic speech, requiring technical knowledge. Voice actors perform live or recorded voice work, emphasizing acting skills and emotional expression. While Ai TTS is automated and scalable, voice actors provide personalized, nuanced performances. Both roles are essential in media and technology industries, but they differ significantly in skills and work environment.
What is AI Text to Speech?
What are some common challenges faced by AI Text-to-Speech specialists, and how are they addressed in a typical work environment?
What are the key skills and qualifications needed to thrive as an AI Text-to-Speech Engineer, and why are they important?

Contractor
Posted yesterday
Job description
Update as of 07/21/26: Hi,
We are specifically seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS), including experience developing and working with modern SOTA architectures and models. It is essential that their TTS background reflects work with the latest technologies and innovations in speech synthesis, demonstrating current industry knowledge and practical implementation experience.
Thank you!
Location: 100% Remote (Anywhere in the U.S.)
Duration: 6-Month Contract
Position Overview
We are seeking a Deep Learning Scientist - Speech Synthesis to support the development of next-generation speech AI technologies. This role focuses on training and optimizing speech models, improving model performance, and solving complex machine learning challenges related to speech applications.
The ideal candidate has strong experience in speech synthesis (Text-to-Speech) or Speech-to-Text, deep learning, and Python development. Success in this role requires the ability to analyze model behavior, diagnose training issues, and improve model performance-not just collect or evaluate data.
Key Responsibilities
- Train and optimize speech synthesis models, including mel spectrogram and vocoder models.
- Analyze training metrics, validation losses, and model performance to identify root causes of model issues and recommend improvements.
- Benchmark and optimize speech models across multiple use cases.
- Improve speech data preparation, augmentation, filtering, and dataset quality.
- Develop and refine high-quality training datasets for speech AI models.
- Measure and characterize model accuracy, quality, and bias.
- Collaborate with cross-functional teams to develop and deliver new speech AI features.
- Participate in software development, design reviews, testing, and code reviews.
- Troubleshoot technical issues and contribute to continuous model improvements.
- Master's degree or Ph.D. in Computer Science, Electrical Engineering, Artificial Intelligence, Applied Mathematics, Linguistics, Computational Linguistics, or a related field (or equivalent experience).
- 3+ years of relevant industry experience.
- Strong Python programming skills.
- Strong understanding of machine learning and deep learning concepts.
- Experience with Text-to-Speech (TTS), Speech Synthesis, or Speech-to-Text (STT) technologies.
- Hands-on experience training deep learning models using PyTorch.
- Ability to analyze training behavior, validation losses, and model performance to troubleshoot and improve machine learning models.
- Knowledge of speech signal processing concepts, including FFT, MFCC, and mel spectrograms.
- Strong understanding of software development fundamentals.
- Experience using version control systems such as Git, Gerrit, or GitLab.
- Excellent communication and collaboration skills.
- Experience with deep learning architectures such as CNNs, RNNs, LSTMs, and Transformers.
- Experience with voice cloning or multilingual speech systems.
- Knowledge of text normalization (TN), inverse text normalization (ITN), or grapheme-to-phoneme (G2P) systems.
- Fluency in one or more languages such as Spanish, Mandarin, German, Japanese, Russian, French, Arabic, Hindi, Korean, Italian, or Portuguese.
- Interest in linguistics, phonetics, and speech technologies.
- Strong C++ programming skills.
- Familiarity with GPU technologies such as CUDA, cuDNN, or TensorRT.
- Experience deploying machine learning models to cloud, data center, or embedded environments.
The ideal candidate is someone who enjoys solving difficult machine learning problems and has hands-on experience training speech models. Beyond building models, we're looking for someone who can investigate why a model is underperforming, analyze validation losses, identify root causes, and improve overall model quality and performance.
Additional Information
- 100% remote position within the United States.
- No specific U.S. time zone requirement.
- This is a contract opportunity.
- Opportunity to contribute to cutting-edge speech AI and deep learning technologies.
About Catapult Solutions Group
Sourced by ZipRecruiter
Industry
Recruiting and staffing services
Company size
201 - 500 Employees
Headquarters location
Plano, TX, US
Year founded
2013