1

Ai Text To Speech Jobs (NOW HIRING)

Build conversational AI systems using LLMs and agent frameworks ... Integrate Speech-to-Text (STT) and Text-to-Speech (TTS) technologies into production applications.

New

Job Summary : Ursus, Inc. is a global leader in AI and high-performance computing, powering ... Text-to-Speech (TTS) systems • Develop WFST and Neural Networks-based Text-Normalization and ...

Innodata builds the high-quality voice and audio datasets that power the world's leading speech AI - text-to-speech, speech recognition, and the new generation of speech-to-speech and conversational ...

Lead Engineer - AI Agent Voice Experience

$104K - $138K/yr

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

Cresta's unified AI platform combines conversational AI agents, real-time human agent augmentation ... text to speech, speech-to-speech models, and production optimization. You will partner closely with ...

Lead Engineer - AI Agent Voice Experience

OR · On-site +1

  • Medical

  • Dental

  • Vision

  • Life

  • Retirement

  • PTO

... text to speech, speech-to-speech models, and production optimization. You will partner closely with ... Improve the quality of voice AI systems through error analysis, data curation, metric design ...

... Text-to-Speech (TTS). • Setting up workflow automations for lead qualification, booking ... Required : • Familiarity with AI workflows • Natural language processing (NLP) pipelines • ...

This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) technology capable of producing ...

next page

Showing results 1-20

Ai Text To Speech information

See salary details

$15

$43

$69

How much do ai text to speech jobs pay per hour?

As of Aug 14, 2026, the average hourly pay for ai text to speech in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is the difference between Ai Text To Speech vs Voice Actor?

AspectAi Text To SpeechVoice Actor
CredentialsNone required, but technical skills helpfulVoice training, acting skills, often a demo reel
Work EnvironmentRemote, software-basedStudio or on-location recording
Industry UsageTechnology, media, customer serviceEntertainment, advertising, narration
Work NatureAutomated voice generation, programmingLive or pre-recorded voice performances

Ai Text To Speech involves using software to convert text into synthetic speech, requiring technical knowledge. Voice actors perform live or recorded voice work, emphasizing acting skills and emotional expression. While Ai TTS is automated and scalable, voice actors provide personalized, nuanced performances. Both roles are essential in media and technology industries, but they differ significantly in skills and work environment.

What is AI Text to Speech?

AI Text to Speech (TTS) is a technology that uses artificial intelligence to convert written text into spoken words. This technology leverages deep learning and neural networks to produce natural-sounding speech that can closely mimic human voices. AI TTS is commonly used in applications such as virtual assistants, accessibility tools for the visually impaired, audiobooks, and automated customer service. It supports multiple languages and can be customized to different voices and accents.

What are some common challenges faced by AI Text-to-Speech specialists, and how are they addressed in a typical work environment?

AI Text-to-Speech specialists often encounter challenges such as ensuring natural-sounding speech synthesis, handling diverse accents, and optimizing for different languages or dialects. Addressing these requires collaborating closely with linguists, data engineers, and software developers to fine-tune models and improve datasets. Regular peer reviews and iterative testing are standard to maintain quality and address edge cases. The work environment is typically cross-functional, fostering open communication to solve problems efficiently.

What are the key skills and qualifications needed to thrive as an AI Text-to-Speech engineer?

To thrive as an AI Text-to-Speech Engineer, you need a strong background in computer science, machine learning, and digital signal processing, typically supported by a relevant degree. Familiarity with tools and frameworks such as TensorFlow, PyTorch, speech synthesis engines, and possibly certification in AI/ML technologies is important. Creativity, problem-solving, and effective collaboration with multidisciplinary teams are crucial soft skills. These abilities enable the development of high-quality, natural-sounding TTS systems that meet user needs and industry standards.
More about Ai Text To Speech jobs

What cities are hiring for Ai Text To Speech jobs?

Cities with the most Ai Text To Speech job openings:

What states have the most Ai Text To Speech jobs?

States with the most job openings for Ai Text To Speech jobs include:

What job categories do people searching Ai Text To Speech jobs look for?

The top searched job categories for Ai Text To Speech jobs are:

Infographic showing various Ai Text To Speech job openings in the United States as of August 2026, with employment types broken down into 50% Full Time, and 50% Part Time. Highlights an 83% In-person, and 17% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Product Manager (AI Speech)

Artificial Analysis

San Francisco, CA • On-site

Full-time

Posted 2 days ago

New


Job description

Job Description - Product Manager (AI Speech)
Location: San Francisco (on-site at our offices)
About Artificial Analysis
Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don't just measure the cutting edge of AI, they are actively shaping the frontier.
Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist.
We are a team of 40+, on track to double by end of year, backed by Nat Friedman (GitHub, Meta), Daniel Gross (SSI, Meta), Andrew Ng (Google Brain, DeepLearning.ai, Amazon), Adam D'Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders.
The Opportunity
Speech is becoming AI's next interface: text to speech, speech to text, voice cloning and real-time voice agents are moving from demos into infrastructure, and our speech benchmarks are how the industry tracks who is winning. We're hiring into our speech pillar to drive that coverage.
You'll build and extend our speech evaluations and arenas, from streaming speech to text and voice cloning to next-generation speech-to-speech intelligence with benchmarks like AgentTalk, and work at a deep technical level with the leading speech AI companies to benchmark their latest models as they launch.
What You'll Do
Benchmark the Speech Frontier: Own coverage across text to speech, speech to text and speech to speech, benchmarking new models and providers as they launch
Design Speech Evaluations: Build and extend the evaluation frameworks, prompt libraries and arenas that reflect how developers and creators actually use speech models, including agentic voice through AgentTalk
Partner with Speech Leaders: Work with the top speech AI companies in the world at a deep technical level to benchmark their models and shape how the industry measures voice
Publish Influential Analysis: Produce the leaderboards, reports and analysis that shape how the industry understands speech AI progress
Drive Product Direction: Shape the roadmap of our speech benchmarking platform together with our engineers and pillar lead
Become AI-Native: Embrace an AI-native workflow, using cutting-edge AI tools to generate leverage in a fast-changing industry and maintain our competitive edge in AI benchmarking
What We're Looking For
You should know speech AI from the inside.
Backgrounds include: product, research or engineering roles at speech AI companies (e.g. ElevenLabs, Cartesia, Inworld, Sesame, Deepgram, AssemblyAI, Rime or similar), voice teams at larger platforms (e.g. OpenAI, Google, Microsoft), or teams building products on text to speech, speech to text or real-time voice.
Required:
• 3+ years of professional experience, including at least 1 year working hands-on with speech AI
• Strong analytical and critical thinking skills
• Proficiency in Python and data analysis
• Hands-on familiarity with modern speech models and their evaluation: quality assessment, word error rate and latency measurement, and preference testing
• Genuine, demonstrable interest and knowledge of Frontier AI. We want people who have informed opinions about where AI is heading, not just people who use AI tools
Why Artificial Analysis?
Shape how AI gets built: The leading AI labs track our benchmarks and use them to guide their development priorities. Your work will directly influence the direction of AI.
Become a world expert in AI: You will evaluate every major model, across every major capability, as they are released. Very few roles offer this breadth of exposure to frontier AI.
Work with the most important players in AI: You'll manage relationships with teams at the leading AI labs and major enterprises as a trusted, independent voice.
Join at a defining moment: We're 40+ people, on track to double by end of year, backed by some of the most connected investors in AI. The people who join now will shape the product, the team, and the strategy as we scale.
Competitive compensation including equity
1