1

Speech Synthesis Jobs (NOW HIRING)

Strong background in one or more of: diffusion models, video generation, generative AI, computer vision, multimodal AI (audio + vision), speech synthesis, or neural rendering. * Experience with ...

Vocal Synthesis - Research in vocal and speech synthesis, along with related areas such as ML-based audio processing and signal processing. Post-Training - Research in post-training techniques for ...

Vocal Synthesis -- Research in vocal and speech synthesis, along with related areas such as ML-based audio processing and signal processing. Post-Training -- Research in post-training techniques for ...

Preferred : • Familiarity with voice cloning, speech synthesis, or related audio ML workflows is a plus. Company : At Uare.ai, we're building the first platform for Individual AI - designed to ...

Preferred : • Familiarity with voice cloning, speech synthesis, or related audio ML workflows is a plus. Company : At Uare.ai, we're building the first platform for Individual AI - designed to ...

Hands-on experience with voice -- conversational AI, speech synthesis, or audio production * Experience migrating NLU systems to LLM and understanding key use cases for RAG functionality improvements

Preferred : • Familiarity with voice cloning, speech synthesis, or related audio ML workflows is a plus. Company : At Uare.ai, we're building the first platform for Individual AI - designed to ...

AI Research Engineer- Speech 1

Redmond, WA · On-site

$229K/yr

Contribute to Speech-to-Speech (S2S) system development, including speech understanding, dialogue management, and speech synthesis components. Research and implement alignment mechanisms between ...

AI Research Engineer- Speech 1

Redmond, WA · On-site

$229K/yr

Contribute to Speech-to-Speech (S2S) system development, including speech understanding, dialogue management, and speech synthesis components. * Research and implement alignment mechanisms between ...

Hands-on experience with voice -- conversational AI, speech synthesis, or audio production * Experience migrating NLU systems to LLM and understanding key use cases for RAG functionality improvements

Agent Experience Designer

New York, NY · On-site

$180K - $220K/yr

Hands-on experience with voice - conversational AI, speech synthesis, or audio production * Experience migrating NLU systems to LLM and understanding key use cases for RAG functionality improvements

Conversation Designer

San Francisco, CA · On-site

$160K - $200K/yr

Hands-on experience with voice -- conversational AI, speech synthesis, or audio production * Experience migrating NLU systems to LLM and understanding key use cases for RAG functionality improvements

Conversation Designer

New York, NY · On-site

$180K - $220K/yr

Hands-on experience with voice - conversational AI, speech synthesis, or audio production * Experience migrating NLU systems to LLM and understanding key use cases for RAG functionality improvements

$160K - $200K/yr

Hands-on experience with voice -- conversational AI, speech synthesis, or audio production * Experience migrating NLU systems to LLM and understanding key use cases for RAG functionality improvements

We have an exciting opportunity for a Software Engineer performing advanced software development in the fields of automatic speech recognition, speech synthesis, Natural Language Processing and ...

Showing results 21-40

Speech Synthesis information

See salary details

$9

$44

$67

How much do speech synthesis jobs pay per hour?

As of Sep 13, 2026, the average hourly pay for speech synthesis in the United States is $44.25, according to ZipRecruiter salary data. Most workers in this role earn between $37.74 and $50.96 per hour, depending on experience, location, and employer.

What is speech synthesis?

Speech synthesis is the artificial production of human speech by computers or other devices. It involves converting written text into spoken words using specialized software and algorithms, often known as text-to-speech (TTS) systems. This technology is used in various applications, such as virtual assistants, accessibility tools for visually impaired users, and automated customer service systems. Advances in machine learning and artificial intelligence have greatly improved the naturalness and clarity of synthetic speech, making it sound more human-like.

What are the key skills and qualifications needed to thrive as a speech synthesis engineer?

To thrive as a Speech Synthesis Engineer, you need a strong background in computer science, linguistics, and signal processing, typically supported by a relevant degree. Expertise with machine learning frameworks (such as TensorFlow or PyTorch), speech synthesis toolkits (like Tacotron or WaveNet), and programming languages (Python or C++) is essential. Strong analytical thinking, creativity, and effective teamwork are standout soft skills for developing natural-sounding and innovative speech technologies. These skills are crucial to advancing accessible, high-quality voice interfaces and applications in diverse industries.

What are some common challenges faced by speech synthesis engineers when developing natural-sounding voices?

Speech synthesis engineers often encounter challenges such as ensuring the generated voice sounds natural and expressive across various contexts and emotions. Achieving accurate pronunciation, intonation, and rhythm requires careful tuning of models and datasets, as well as extensive testing across diverse languages and accents. Collaboration with linguists, data scientists, and voice talent is common to refine vocal characteristics and address edge cases. Staying updated with advances in deep learning and neural network architectures is also important for continuous improvement.

What is the difference between Speech Synthesis vs Speech Recognition?

AspectSpeech SynthesisSpeech Recognition
Required CredentialsComputer Science, Linguistics, Audio EngineeringComputer Science, Linguistics, Signal Processing
Work EnvironmentSoftware development, AI labs, tech companiesCall centers, voice assistant companies, research labs
Industry UsageGenerating spoken output from textConverting spoken input into text

Speech Synthesis involves creating artificial speech from text, enabling applications like text-to-speech systems. Speech Recognition focuses on converting spoken language into written text, used in voice assistants and transcription services. While both involve audio processing and require similar technical skills, they serve opposite functions in voice technology.

What does speech synthesis do?

Speech synthesis is a technology used in the Speech Synthesis job to convert written text into spoken words using algorithms and digital voices. It involves designing and improving voice quality, intonation, and clarity, often utilizing tools like text analysis and signal processing. This process enables applications such as virtual assistants, screen readers, and language learning tools.

What cities are hiring for Speech Synthesis jobs?

Cities with the most Speech Synthesis job openings:

What states have the most Speech Synthesis jobs?

States with the most job openings for Speech Synthesis jobs include:

What job categories do people searching Speech Synthesis jobs look for?

The top searched job categories for Speech Synthesis jobs are:

What other helpful pages are available for Speech Synthesis?

Other pages related to Speech Synthesis:

Infographic showing various Speech Synthesis job openings in the United States as of September 2026, with employment types broken down into 5% As Needed, 60% Full Time, 22% Part Time, and 13% Contract. Highlights an 93% Physical, 1% Hybrid, and 6% Remote job distribution, with an average salary of $92,039 per year, or $44.2 per hour.

Information Technology_USA - USA_Developer

Jacksonville, FL • On-site

Real Soft, Inc.
IT Services • 501 - 1,000 employees

Contractor

This job post has expired 1 day ago. Applications are no longer accepted.


Job description

**Please strictly adhere to the following resume naming convention:
ALL CAPS, NO SPACES BETWEEN UNDERSCORES
PTN_US_GBAMSREQID_CandidateBeelineID
Example: PTN_US_9999999_SKIPJOHNSON0413
: -
MSP Owner: Michelle Lee
Location: Hartford, CT
Duration: 6 months
skill id: 10666754
Required Qualifications
5+ years of overall software delivery experience.
3+ years of chatbot development experience using IBM WatsonX Orchestrate or similar conversational AI platforms, such as:
Google Dialogflow CX
LivePerson
Amazon Lex
Kore.ai
3+ years of experience training and improving intent recognition, including curating and refining training datasets.
3+ years of experience applying Conversational AI best practices, including:
Natural Language Processing (NLP)
Training data design and optimization
1+ year of experience with:
CI/CD pipelines
Git version control
Unit testing
Source code management
1+ year of knowledge in cloud development and deployment principles.
1+ year of backend (server-side) development experience using:
Java
Node.js
Python
1+ year of multi-language development experience using:
JavaScript
TypeScript
1+ year of experience working in an Agile environment, with understanding of:
DevOps practices
End-to-end software development lifecycles
Business requirements translation
Preferred Qualifications
Prior experience with:
IBM Watson Speech-to-Text
IBM Watson Discovery
Similar voice processing or information retrieval technologies
Familiarity with SSML (Speech Synthesis Markup Language) for text-to-speech (TTS) refinement.
Experience integrating unstructured data sources-such as document repositories-into conversational agents using semantic embeddings or advanced search techniques.
Skills: Digital : Python ~ Digital : Natural Language Processing (NLP)~Digital : Node.js ~ Digital : DevOps ~ Foundation : JavaScript
Experience Required: 10 & Above, Project Code :