1

Ai Text To Speech Jobs (NOW HIRING)

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI. Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech ...

Lead Engineer - AI Agent Voice Experience

$104K - $138K/yr

Cresta's unified AI platform combines conversational AI agents, real-time human agent augmentation ... text to speech, speech-to-speech models, and production optimization. You will partner closely with ...

... text to speech, speech-to-speech models, and production optimization. You will partner closely with ... Improve the quality of voice AI systems through error analysis, data curation, metric design ...

FDE with Amazon Connect

Saint Louis, MO ยท On-site

$73K - $170K/yr

Speech-to-speech (S2S) model integration for real-time voice AI * Text-to-speech (TTS) and speech-to-text (STT) pipeline design and optimization * Latency optimization for conversational voice ...

Showing results 21-40

Ai Text To Speech information

See salary details

$15

$43

$69

How much do ai text to speech jobs pay per hour?

As of Sep 3, 2026, the average hourly pay for ai text to speech in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is AI Text to Speech?

AI Text to Speech (TTS) is a technology that uses artificial intelligence to convert written text into spoken words. This technology leverages deep learning and neural networks to produce natural-sounding speech that can closely mimic human voices. AI TTS is commonly used in applications such as virtual assistants, accessibility tools for the visually impaired, audiobooks, and automated customer service. It supports multiple languages and can be customized to different voices and accents.

What are the key skills and qualifications needed to thrive as an AI Text-to-Speech engineer?

To thrive as an AI Text-to-Speech Engineer, you need a strong background in computer science, machine learning, and digital signal processing, typically supported by a relevant degree. Familiarity with tools and frameworks such as TensorFlow, PyTorch, speech synthesis engines, and possibly certification in AI/ML technologies is important. Creativity, problem-solving, and effective collaboration with multidisciplinary teams are crucial soft skills. These abilities enable the development of high-quality, natural-sounding TTS systems that meet user needs and industry standards.

What are some common challenges faced by AI Text-to-Speech specialists, and how are they addressed in a typical work environment?

AI Text-to-Speech specialists often encounter challenges such as ensuring natural-sounding speech synthesis, handling diverse accents, and optimizing for different languages or dialects. Addressing these requires collaborating closely with linguists, data engineers, and software developers to fine-tune models and improve datasets. Regular peer reviews and iterative testing are standard to maintain quality and address edge cases. The work environment is typically cross-functional, fostering open communication to solve problems efficiently.

What is the difference between Ai Text To Speech vs Voice Actor?

AspectAi Text To SpeechVoice Actor
CredentialsNone required, but technical skills helpfulVoice training, acting skills, often a demo reel
Work EnvironmentRemote, software-basedStudio or on-location recording
Industry UsageTechnology, media, customer serviceEntertainment, advertising, narration
Work NatureAutomated voice generation, programmingLive or pre-recorded voice performances

Ai Text To Speech involves using software to convert text into synthetic speech, requiring technical knowledge. Voice actors perform live or recorded voice work, emphasizing acting skills and emotional expression. While Ai TTS is automated and scalable, voice actors provide personalized, nuanced performances. Both roles are essential in media and technology industries, but they differ significantly in skills and work environment.

More about Ai Text To Speech jobs

What cities are hiring for Ai Text To Speech jobs?

Cities with the most Ai Text To Speech job openings:

What states have the most Ai Text To Speech jobs?

States with the most job openings for Ai Text To Speech jobs include:

Infographic showing various Ai Text To Speech job openings in the United States as of August 2026, with employment types broken down into 76% Full Time, 20% Part Time, and 4% Contract. Highlights an 66% Physical, 4% Hybrid, and 30% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Deep Learning Scientist, Speech Synthesis

Vailexa

Santa Clara, CA โ€ข On-site

Full-time

Re-posted 14 days ago


Job description

Build More Than Just a Career. Build Your Future.

At Vailexa, we're not just hiring — we're building thinkers, creators, and future leaders.

We believe in giving people the space to grow, the freedom to think, and the opportunity to create real impact from day one. If you're someone who wants to learn fast, take ownership, and grow beyond limits, you'll feel right at home here.

Key Responsibilities

  • Train speech synthesis mel spectrogram and vocoder models
  • Measure and benchmark model performance across use cases
  • Maintain and enhance text to speech evaluation systems
  • Analyze model accuracy and bias and recommend improvements
  • Improve processes related to speech data preparation, augmentation, and filtering
  • Develop and refine training datasets for speech models
  • Characterize performance and quality metrics across different platforms
  • Collaborate with cross functional teams to deliver new product features
  • Participate in code development, design reviews, and test planning
  • Identify issues, propose solutions, and contribute to continuous innovation

Required Qualifications

  • Master's degree or PhD in Computer Science, Electrical Engineering, Artificial Intelligence, Applied Mathematics, Linguistics, or Computational Linguistics or equivalent experience
  • Minimum of 5 years of relevant experience
  • Strong programming skills in Python
  • Solid understanding of programming fundamentals and software design
  • Deep knowledge of machine learning and deep learning techniques including CNN, RNN, LSTM, and Transformers
  • Experience applying deep learning to speech synthesis, large language models, and speech to speech translation
  • Hands on experience with speech technologies such as speech synthesis and voice cloning
  • Experience training speech models
  • Proficiency with PyTorch deep learning frameworks
  • Knowledge of speech signal processing techniques including FFT, MFCC, and mel spectrograms
  • Familiarity with version control tools such as Git, Gerrit, or GitLab
  • Strong collaboration and communication skills in a matrixed environment

Preferred Qualifications

  • Fluency in one or more languages such as Spanish, Mandarin, German, Japanese, Russian, French, Arabic, Hindi, Korean, Italian, or Portuguese
  • Experience with multilingual or code switched text to speech systems
  • Experience with voice cloning and cross lingual voice cloning
  • Knowledge of text normalization and inverse text normalization using neural networks or WFST
  • Experience working with grapheme to phoneme systems for multiple languages
  • Interest in linguistics, phonetics, and language technologies
  • Strong C plus plus programming skills
  • Familiarity with GPU technologies such as CUDA, cuDNN, or TensorRT
  • Experience deploying machine learning models to cloud, data center, or embedded systems

Ready to take the next step?

If you're excited about this role and ready to grow with a team that values ambition, ideas, and impact — we'd love to hear from you.

???? Apply now and start building your journey with Vailexa.