1

Ai Text To Speech Jobs (NOW HIRING)

You will lead the design and delivery of end-to-end voice AI solutions, combining large language models with speech technologies such as speech-to-text, text-to-speech, and real-time streaming audio ...

ML Researcher, Speech

San Francisco, CA · On-site

$200K - $250K/yr

... to-speech conversational AI model that understands and responds like a human, in real time. You'll work across the core building blocks of that roadmap, such as: speech-to-text, text-to-speech ...

Innodata builds the high-quality voice and audio datasets that power the world's leading speech AI - text-to-speech, speech recognition, and the new generation of speech-to-speech and conversational ...

Lead Engineer - AI Agent Voice Experience

$104K - $138K/yr

Cresta's unified AI platform combines conversational AI agents, real-time human agent augmentation ... text to speech, speech-to-speech models, and production optimization. You will partner closely with ...

... text to speech, speech-to-speech models, and production optimization. You will partner closely with ... Improve the quality of voice AI systems through error analysis, data curation, metric design ...

... Text-to-Speech (TTS). • Setting up workflow automations for lead qualification, booking ... Required : • Familiarity with AI workflows • Natural language processing (NLP) pipelines • ...

This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) technology capable of producing ...

next page

Showing results 1-20

Ai Text To Speech information

See salary details

$15

$43

$69

How much do ai text to speech jobs pay per hour?

As of Jul 26, 2026, the average hourly pay for ai text to speech in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is the difference between Ai Text To Speech vs Voice Actor?

AspectAi Text To SpeechVoice Actor
CredentialsNone required, but technical skills helpfulVoice training, acting skills, often a demo reel
Work EnvironmentRemote, software-basedStudio or on-location recording
Industry UsageTechnology, media, customer serviceEntertainment, advertising, narration
Work NatureAutomated voice generation, programmingLive or pre-recorded voice performances

Ai Text To Speech involves using software to convert text into synthetic speech, requiring technical knowledge. Voice actors perform live or recorded voice work, emphasizing acting skills and emotional expression. While Ai TTS is automated and scalable, voice actors provide personalized, nuanced performances. Both roles are essential in media and technology industries, but they differ significantly in skills and work environment.

What is AI Text to Speech?

AI Text to Speech (TTS) is a technology that uses artificial intelligence to convert written text into spoken words. This technology leverages deep learning and neural networks to produce natural-sounding speech that can closely mimic human voices. AI TTS is commonly used in applications such as virtual assistants, accessibility tools for the visually impaired, audiobooks, and automated customer service. It supports multiple languages and can be customized to different voices and accents.

What are some common challenges faced by AI Text-to-Speech specialists, and how are they addressed in a typical work environment?

AI Text-to-Speech specialists often encounter challenges such as ensuring natural-sounding speech synthesis, handling diverse accents, and optimizing for different languages or dialects. Addressing these requires collaborating closely with linguists, data engineers, and software developers to fine-tune models and improve datasets. Regular peer reviews and iterative testing are standard to maintain quality and address edge cases. The work environment is typically cross-functional, fostering open communication to solve problems efficiently.

What are the key skills and qualifications needed to thrive as an AI Text-to-Speech Engineer, and why are they important?

To thrive as an AI Text-to-Speech Engineer, you need a strong background in computer science, machine learning, and digital signal processing, typically supported by a relevant degree. Familiarity with tools and frameworks such as TensorFlow, PyTorch, speech synthesis engines, and possibly certification in AI/ML technologies is important. Creativity, problem-solving, and effective collaboration with multidisciplinary teams are crucial soft skills. These abilities enable the development of high-quality, natural-sounding TTS systems that meet user needs and industry standards.
More about Ai Text To Speech jobs
What cities are hiring for Ai Text To Speech jobs? Cities with the most Ai Text To Speech job openings:
What states have the most Ai Text To Speech jobs? States with the most job openings for Ai Text To Speech jobs include:
Infographic showing various Ai Text To Speech job openings in the United States as of July 2026, with employment types broken down into 73% Full Time, 24% Part Time, and 3% Contract. Highlights an 65% Physical, 3% Hybrid, and 32% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.
Deep Learning Scientist, Speech Synthesis

Deep Learning Scientist, Speech Synthesis

Catapult Solutions Group

Santa Clara, CA • On-site, Remote

Contractor

Posted yesterday


Job description

Deep Learning Scientist - Speech Synthesis
Update as of 07/21/26: Hi,
We are specifically seeking candidates with deep, hands-on expertise in Text-to-Speech (TTS), including experience developing and working with modern SOTA architectures and models. It is essential that their TTS background reflects work with the latest technologies and innovations in speech synthesis, demonstrating current industry knowledge and practical implementation experience.
Thank you!
Location: 100% Remote (Anywhere in the U.S.)
Duration: 6-Month Contract
Position Overview
We are seeking a Deep Learning Scientist - Speech Synthesis to support the development of next-generation speech AI technologies. This role focuses on training and optimizing speech models, improving model performance, and solving complex machine learning challenges related to speech applications.
The ideal candidate has strong experience in speech synthesis (Text-to-Speech) or Speech-to-Text, deep learning, and Python development. Success in this role requires the ability to analyze model behavior, diagnose training issues, and improve model performance-not just collect or evaluate data.
Key Responsibilities
  • Train and optimize speech synthesis models, including mel spectrogram and vocoder models.
  • Analyze training metrics, validation losses, and model performance to identify root causes of model issues and recommend improvements.
  • Benchmark and optimize speech models across multiple use cases.
  • Improve speech data preparation, augmentation, filtering, and dataset quality.
  • Develop and refine high-quality training datasets for speech AI models.
  • Measure and characterize model accuracy, quality, and bias.
  • Collaborate with cross-functional teams to develop and deliver new speech AI features.
  • Participate in software development, design reviews, testing, and code reviews.
  • Troubleshoot technical issues and contribute to continuous model improvements.
Required Qualifications
  • Master's degree or Ph.D. in Computer Science, Electrical Engineering, Artificial Intelligence, Applied Mathematics, Linguistics, Computational Linguistics, or a related field (or equivalent experience).
  • 3+ years of relevant industry experience.
  • Strong Python programming skills.
  • Strong understanding of machine learning and deep learning concepts.
  • Experience with Text-to-Speech (TTS), Speech Synthesis, or Speech-to-Text (STT) technologies.
  • Hands-on experience training deep learning models using PyTorch.
  • Ability to analyze training behavior, validation losses, and model performance to troubleshoot and improve machine learning models.
  • Knowledge of speech signal processing concepts, including FFT, MFCC, and mel spectrograms.
  • Strong understanding of software development fundamentals.
  • Experience using version control systems such as Git, Gerrit, or GitLab.
  • Excellent communication and collaboration skills.
Preferred Qualifications
  • Experience with deep learning architectures such as CNNs, RNNs, LSTMs, and Transformers.
  • Experience with voice cloning or multilingual speech systems.
  • Knowledge of text normalization (TN), inverse text normalization (ITN), or grapheme-to-phoneme (G2P) systems.
  • Fluency in one or more languages such as Spanish, Mandarin, German, Japanese, Russian, French, Arabic, Hindi, Korean, Italian, or Portuguese.
  • Interest in linguistics, phonetics, and speech technologies.
  • Strong C++ programming skills.
  • Familiarity with GPU technologies such as CUDA, cuDNN, or TensorRT.
  • Experience deploying machine learning models to cloud, data center, or embedded environments.
What We're Looking For
The ideal candidate is someone who enjoys solving difficult machine learning problems and has hands-on experience training speech models. Beyond building models, we're looking for someone who can investigate why a model is underperforming, analyze validation losses, identify root causes, and improve overall model quality and performance.
Additional Information
  • 100% remote position within the United States.
  • No specific U.S. time zone requirement.
  • This is a contract opportunity.
  • Opportunity to contribute to cutting-edge speech AI and deep learning technologies.