No PHD is need just focus on seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS), including experience developing and working with modern SOTA architectures and models. It is ...
No PHD is need just focus on seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS), including experience developing and working with modern SOTA architectures and models. It is ...
Maintain and enhance text to speech evaluation systems * Analyze model accuracy and bias and recommend improvements * Improve processes related to speech data preparation, augmentation, and filtering
Quick apply
Maintain and enhance text to speech evaluation systems * Analyze model accuracy and bias and recommend improvements * Improve processes related to speech data preparation, augmentation, and filtering
AI Research Initiative For Text-To-Speech Systems This role is for one of our clients. Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text ...
AI Research Initiative For Text-To-Speech Systems This role is for one of our clients. Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text ...
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT) and text-to-speech (TTS). They are seeking a highly ...
Deepgram is the leading platform underpinning the emerging trillion-dollar Voice AI economy, providing real-time APIs for speech-to-text (STT) and text-to-speech (TTS). They are seeking a highly ...
Responsibilities : • Build speech training and evaluation data sets for Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems • Develop WFST and Neural Networks-based Text ...
Responsibilities : • Build speech training and evaluation data sets for Automatic Speech Recognition (ASR) and Text-to-Speech (TTS) systems • Develop WFST and Neural Networks-based Text ...
ML Researcher, Speech
San Francisco, CA · On-site
$200K - $250K/yr
You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech, neural audio codecs, and getting LLMs to understand and reason over audio directly. You'll take ...
ML Researcher, Speech
San Francisco, CA · On-site
$200K - $250K/yr
You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech, neural audio codecs, and getting LLMs to understand and reason over audio directly. You'll take ...
Expertise in speech-to-text, text-to-speech, and designing natural, human-like voice conversations. • NLP & Machine Learning: Strong background in Natural Language Processing, intent recognition ...
Expertise in speech-to-text, text-to-speech, and designing natural, human-like voice conversations. • NLP & Machine Learning: Strong background in Natural Language Processing, intent recognition ...
Whether text-to-text, text-to-speech, speech-to-text, or speech-to-speech, machine translation can help overlooked communities finally be understood in the world. HLT will bring critical educational ...
Whether text-to-text, text-to-speech, speech-to-text, or speech-to-speech, machine translation can help overlooked communities finally be understood in the world. HLT will bring critical educational ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) systems capable of producing ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) systems capable of producing ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) systems capable of producing ...
This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) systems capable of producing ...
Information Technology_USA - USA_Engineer
Jacksonville, FL · On-site
$49 - $67/hr
... text/text-to-speech. Role Descriptions: 4+ years of commercial software development experience.Design and implement scalable CCaaS and IVA solutions leveraging leading Cloud and enterprise ...
Information Technology_USA - USA_Engineer
Jacksonville, FL · On-site
$49 - $67/hr
... text/text-to-speech. Role Descriptions: 4+ years of commercial software development experience.Design and implement scalable CCaaS and IVA solutions leveraging leading Cloud and enterprise ...
... Text-to-Speech), and NLP/LLM pipelines. • Create frameworks for conversational flows, prompt engineering, retrieval-augmented generation (RAG), and context management. Solution Development • ...
... Text-to-Speech), and NLP/LLM pipelines. • Create frameworks for conversational flows, prompt engineering, retrieval-augmented generation (RAG), and context management. Solution Development • ...
Expertise and intuition for training models in the audio domain, including text-to-speech, ASR, speech-to-speech, speech-emotion-recognition, or other models * Experience in training audio ...
Expertise and intuition for training models in the audio domain, including text-to-speech, ASR, speech-to-speech, speech-emotion-recognition, or other models * Experience in training audio ...
Expertise and intuition for training models in the audio domain, including text-to-speech, ASR, speech-to-speech, speech-emotion-recognition, or other models * Experience in training audio ...
Quick apply
Expertise and intuition for training models in the audio domain, including text-to-speech, ASR, speech-to-speech, speech-emotion-recognition, or other models * Experience in training audio ...
Speech recognition & Text to Speech * Web & Cloud technologies Skills: * IVR, Genesys, Java, Telephony, SQL, UNIX, Windows, Oracle
Speech recognition & Text to Speech * Web & Cloud technologies Skills: * IVR, Genesys, Java, Telephony, SQL, UNIX, Windows, Oracle
Software Engineer, Platform - Gainesville, FL, USA
Gainesville, TX · On-site
$140 - $200/hr
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App.Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Software Engineer, Platform - Gainesville, FL, USA
Gainesville, TX · On-site
$140 - $200/hr
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App.Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Deepgram is the leading voice AI platform for developers building speech-to-text and text-to-speech offerings. They are seeking a Software Engineer to join their new business unit focused on ...
Deepgram is the leading voice AI platform for developers building speech-to-text and text-to-speech offerings. They are seeking a Software Engineer to join their new business unit focused on ...
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App. Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App. Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App. Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App. Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App. Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Speechify's text-to-speech reading products include its iOS app, Android App, Mac App, Chrome Extension, and Web App. Google recently named Speechify the Chrome Extension of the Year and Apple named ...
Text To Speech information
What is a text to speech job?
What are some common challenges faced by professionals working in text to speech development roles?
What is the difference between Text To Speech vs Voice Actor?
| Aspect | Text To Speech | Voice Actor |
|---|---|---|
| Required Credentials | None or basic audio editing skills | Voice training, acting skills, often professional demos |
| Work Environment | Software, digital platforms, remote | Recording studios, on-location, remote |
| Industry Usage | Automation, AI, tech companies | Media, entertainment, advertising |
| Search & Comparison Intent | Automated voice solutions, TTS technology | Voice acting, narration, character voices |
Text To Speech involves using software to convert written text into spoken words, primarily for automation and digital applications. Voice Actors, on the other hand, provide human voice recordings for media, entertainment, and advertising. While TTS is tech-driven and often used in AI and accessibility tools, Voice Actors bring emotional nuance and personality to their performances. Both roles are essential in their respective industries, but they differ significantly in skills, environment, and purpose.
What are the key skills and qualifications needed to thrive as a text to speech engineer?

Contractor
Re-posted 16 days ago
Job description
Update on 07/22/26:
No PHD is need just focus on seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS), including experience developing and working with modern SOTA architectures and models. It is essential that their TTS background reflects work with the latest technologies and innovations in speech synthesis, demonstrating current industry knowledge and practical implementation experience.
Update as of 07/21/26: Hi,
We are specifically seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS), including experience developing and working with modern SOTA architectures and models. It is essential that their TTS background reflects work with the latest technologies and innovations in speech synthesis, demonstrating current industry knowledge and practical implementation experience.
Thank you!
Location: 100% Remote (Anywhere in the U.S.)
Duration: 6-Month Contract
Position Overview
We are seeking a Deep Learning Scientist - Speech Synthesis to support the development of next-generation speech AI technologies. This role focuses on training and optimizing speech models, improving model performance, and solving complex machine learning challenges related to speech applications.
The ideal candidate has strong experience in speech synthesis (Text-to-Speech) or Speech-to-Text, deep learning, and Python development. Success in this role requires the ability to analyze model behavior, diagnose training issues, and improve model performance-not just collect or evaluate data.
Key Responsibilities
- Train and optimize speech synthesis models, including mel spectrogram and vocoder models.
- Analyze training metrics, validation losses, and model performance to identify root causes of model issues and recommend improvements.
- Benchmark and optimize speech models across multiple use cases.
- Improve speech data preparation, augmentation, filtering, and dataset quality.
- Develop and refine high-quality training datasets for speech AI models.
- Measure and characterize model accuracy, quality, and bias.
- Collaborate with cross-functional teams to develop and deliver new speech AI features.
- Participate in software development, design reviews, testing, and code reviews.
- Troubleshoot technical issues and contribute to continuous model improvements.
- Master's degree or Ph.D. in Computer Science, Electrical Engineering, Artificial Intelligence, Applied Mathematics, Linguistics, Computational Linguistics, or a related field (or equivalent experience).
- 3+ years of relevant industry experience.
- Strong Python programming skills.
- Strong understanding of machine learning and deep learning concepts.
- Experience with Text-to-Speech (TTS), Speech Synthesis, or Speech-to-Text (STT) technologies.
- Hands-on experience training deep learning models using PyTorch.
- Ability to analyze training behavior, validation losses, and model performance to troubleshoot and improve machine learning models.
- Knowledge of speech signal processing concepts, including FFT, MFCC, and mel spectrograms.
- Strong understanding of software development fundamentals.
- Experience using version control systems such as Git, Gerrit, or GitLab.
- Excellent communication and collaboration skills.
- Experience with deep learning architectures such as CNNs, RNNs, LSTMs, and Transformers.
- Experience with voice cloning or multilingual speech systems.
- Knowledge of text normalization (TN), inverse text normalization (ITN), or grapheme-to-phoneme (G2P) systems.
- Fluency in one or more languages such as Spanish, Mandarin, German, Japanese, Russian, French, Arabic, Hindi, Korean, Italian, or Portuguese.
- Interest in linguistics, phonetics, and speech technologies.
- Strong C++ programming skills.
- Familiarity with GPU technologies such as CUDA, cuDNN, or TensorRT.
- Experience deploying machine learning models to cloud, data center, or embedded environments.
The ideal candidate is someone who enjoys solving difficult machine learning problems and has hands-on experience training speech models. Beyond building models, we're looking for someone who can investigate why a model is underperforming, analyze validation losses, identify root causes, and improve overall model quality and performance.
Additional Information
- 100% remote position within the United States.
- No specific U.S. time zone requirement.
- This is a contract opportunity.
- Opportunity to contribute to cutting-edge speech AI and deep learning technologies.
About Catapult Solutions Group
Sourced by ZipRecruiter
Industry
Recruiting and staffing services
Company size
201 - 500 Employees
Headquarters location
Plano, TX, US
Year founded
2013