1

Speech Synthesis Jobs (NOW HIRING)

Familiarity with voice cloning, speech synthesis, or related audio ML workflows is a plus. * Excited about building 0-to-1 products in a fast-moving environment. The expected base compensation for ...

Familiarity with voice cloning, speech synthesis, or related audio ML workflows is a plus. * Excited about building 0-to-1 products in a fast-moving environment. The expected base compensation for ...

Deep expertise in at least one of: speech synthesis, ASR, audio modeling, language modeling, or multimodal learning * Fluency in the math, probability, optimization, linear algebra, and the ability ...

Deep expertise in at least one of: speech synthesis, ASR, audio modeling, language modeling, or multimodal learning * Fluency in the math, probability, optimization, linear algebra, and the ability ...

Speech Therapist, Part time

Zion, IL · On-site

$38.25 - $52.25/hr

Speech Therapist, Part Time About City of Hope, City of Hope's mission is to make hope a reality ... synthetic insulin and monoclonal antibodies. With an independent, National Cancer Institute ...

Speech Therapist, Part time

Zion, IL · On-site

$38.25 - $52.25/hr

Speech Therapist, Part time About City of Hope, City of Hope's mission is to make hope a reality ... synthetic insulin and monoclonal antibodies. With an independent, National Cancer Institute ...

Showing results 41-60

Speech Synthesis information

See salary details

$9

$44

$67

How much do speech synthesis jobs pay per hour?

As of Sep 13, 2026, the average hourly pay for speech synthesis in the United States is $44.25, according to ZipRecruiter salary data. Most workers in this role earn between $37.74 and $50.96 per hour, depending on experience, location, and employer.

What is speech synthesis?

Speech synthesis is the artificial production of human speech by computers or other devices. It involves converting written text into spoken words using specialized software and algorithms, often known as text-to-speech (TTS) systems. This technology is used in various applications, such as virtual assistants, accessibility tools for visually impaired users, and automated customer service systems. Advances in machine learning and artificial intelligence have greatly improved the naturalness and clarity of synthetic speech, making it sound more human-like.

What are the key skills and qualifications needed to thrive as a speech synthesis engineer?

To thrive as a Speech Synthesis Engineer, you need a strong background in computer science, linguistics, and signal processing, typically supported by a relevant degree. Expertise with machine learning frameworks (such as TensorFlow or PyTorch), speech synthesis toolkits (like Tacotron or WaveNet), and programming languages (Python or C++) is essential. Strong analytical thinking, creativity, and effective teamwork are standout soft skills for developing natural-sounding and innovative speech technologies. These skills are crucial to advancing accessible, high-quality voice interfaces and applications in diverse industries.

What are some common challenges faced by speech synthesis engineers when developing natural-sounding voices?

Speech synthesis engineers often encounter challenges such as ensuring the generated voice sounds natural and expressive across various contexts and emotions. Achieving accurate pronunciation, intonation, and rhythm requires careful tuning of models and datasets, as well as extensive testing across diverse languages and accents. Collaboration with linguists, data scientists, and voice talent is common to refine vocal characteristics and address edge cases. Staying updated with advances in deep learning and neural network architectures is also important for continuous improvement.

What is the difference between Speech Synthesis vs Speech Recognition?

AspectSpeech SynthesisSpeech Recognition
Required CredentialsComputer Science, Linguistics, Audio EngineeringComputer Science, Linguistics, Signal Processing
Work EnvironmentSoftware development, AI labs, tech companiesCall centers, voice assistant companies, research labs
Industry UsageGenerating spoken output from textConverting spoken input into text

Speech Synthesis involves creating artificial speech from text, enabling applications like text-to-speech systems. Speech Recognition focuses on converting spoken language into written text, used in voice assistants and transcription services. While both involve audio processing and require similar technical skills, they serve opposite functions in voice technology.

What does speech synthesis do?

Speech synthesis is a technology used in the Speech Synthesis job to convert written text into spoken words using algorithms and digital voices. It involves designing and improving voice quality, intonation, and clarity, often utilizing tools like text analysis and signal processing. This process enables applications such as virtual assistants, screen readers, and language learning tools.

What cities are hiring for Speech Synthesis jobs?

Cities with the most Speech Synthesis job openings:

What states have the most Speech Synthesis jobs?

States with the most job openings for Speech Synthesis jobs include:

What job categories do people searching Speech Synthesis jobs look for?

The top searched job categories for Speech Synthesis jobs are:

What other helpful pages are available for Speech Synthesis?

Other pages related to Speech Synthesis:

Infographic showing various Speech Synthesis job openings in the United States as of September 2026, with employment types broken down into 5% As Needed, 60% Full Time, 22% Part Time, 1% Temporary, and 12% Contract. Highlights an 93% Physical, 1% Hybrid, and 6% Remote job distribution, with an average salary of $92,039 per year, or $44.2 per hour.

Software Engineer (Early Career)

Los Altos, CA

Full-time

Re-posted 8 days ago


Job description

Uare.ai, founded by Robert LoCascio (former CEO of LivePerson for 28 years), is an AI startup launched in May 2024 with a mission to empower people to do more with their memories. Uare.ai creates AI-driven personal digital twins, enabling users to preserve, share, and interact with their knowledge and stories in groundbreaking ways.

At its core is the proprietary Human Life Model (HLM), designed to process unstructured conversations and story data sets without training client data into traditional LLMs. Instead, Uare.ai leverages 3rd-party LLMs for conversations, with HLM providing structure and reasoning.

While initially focused on personal legacy and AI immortality, Uare.ai' platform has broad B2C and B2B applications. The company has gained significant traction, featured in over 30 outlets including NBC, NPR, The New York Post, and CNET.

 About the Role:

We are seeking a highly skilled Machine Learning Engineer to join an early-stage, mission-driven startup that's building cutting-edge AI-powered products for real-world human interaction. You will be responsible for designing and implementing scalable machine learning pipelines, integrating foundation models, and fine-tuning retrieval systems that deliver high-quality, context-aware responses.

What You'll Do:

  • Design and implement machine learning workflows that combine structured and unstructured data.
  • Work with state-of-the-art language models and voice technologies.
  • Build and optimize retrieval-augmented generation (RAG) systems.
  • Collaborate closely with product and engineering teams to turn early prototypes into production-grade systems.
  • Help shape the future of AI-powered human interaction products. 

What We're Looking For:

  • Experience in machine learning, NLP, or AI.
  • Strong experience with LLMs, RAG pipelines, vector databases, and prompt engineering.
  • Hands-on experience with cloud-based AI platforms (AWS, Azure, or GCP).
  • Familiarity with voice cloning, speech synthesis, or related audio ML workflows is a plus.
  • Excited about building 0-to-1 products in a fast-moving environment.