1

Ai Text To Speech Jobs (NOW HIRING)

Senior AI Engineer - Voice

San Francisco, CA ยท On-site

$123K - $169K/yr

Build and operate the real-time stack - telephony, streaming speech-to-text, LLM orchestration ... You've debugged AI systems in production and can trace a bad outcome through audio, transcript, and ...

Senior Machine Learning Engineer

Sandy, UT ยท Hybrid

$99K - $136K/yr

Stay informed of advances in speech AI, including transcription, text-to-speech, and speech-to-speech technologies. Have you got what it takes? * MS in computer science, electrical engineering ...

Showing results 41-60

Ai Text To Speech information

See salary details

$15

$43

$69

How much do ai text to speech jobs pay per hour?

As of Aug 15, 2026, the average hourly pay for ai text to speech in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is the difference between Ai Text To Speech vs Voice Actor?

AspectAi Text To SpeechVoice Actor
CredentialsNone required, but technical skills helpfulVoice training, acting skills, often a demo reel
Work EnvironmentRemote, software-basedStudio or on-location recording
Industry UsageTechnology, media, customer serviceEntertainment, advertising, narration
Work NatureAutomated voice generation, programmingLive or pre-recorded voice performances

Ai Text To Speech involves using software to convert text into synthetic speech, requiring technical knowledge. Voice actors perform live or recorded voice work, emphasizing acting skills and emotional expression. While Ai TTS is automated and scalable, voice actors provide personalized, nuanced performances. Both roles are essential in media and technology industries, but they differ significantly in skills and work environment.

What is AI Text to Speech?

AI Text to Speech (TTS) is a technology that uses artificial intelligence to convert written text into spoken words. This technology leverages deep learning and neural networks to produce natural-sounding speech that can closely mimic human voices. AI TTS is commonly used in applications such as virtual assistants, accessibility tools for the visually impaired, audiobooks, and automated customer service. It supports multiple languages and can be customized to different voices and accents.

What are some common challenges faced by AI Text-to-Speech specialists, and how are they addressed in a typical work environment?

AI Text-to-Speech specialists often encounter challenges such as ensuring natural-sounding speech synthesis, handling diverse accents, and optimizing for different languages or dialects. Addressing these requires collaborating closely with linguists, data engineers, and software developers to fine-tune models and improve datasets. Regular peer reviews and iterative testing are standard to maintain quality and address edge cases. The work environment is typically cross-functional, fostering open communication to solve problems efficiently.

What are the key skills and qualifications needed to thrive as an AI Text-to-Speech engineer?

To thrive as an AI Text-to-Speech Engineer, you need a strong background in computer science, machine learning, and digital signal processing, typically supported by a relevant degree. Familiarity with tools and frameworks such as TensorFlow, PyTorch, speech synthesis engines, and possibly certification in AI/ML technologies is important. Creativity, problem-solving, and effective collaboration with multidisciplinary teams are crucial soft skills. These abilities enable the development of high-quality, natural-sounding TTS systems that meet user needs and industry standards.
More about Ai Text To Speech jobs

What cities are hiring for Ai Text To Speech jobs?

Cities with the most Ai Text To Speech job openings:

What states have the most Ai Text To Speech jobs?

States with the most job openings for Ai Text To Speech jobs include:

What job categories do people searching Ai Text To Speech jobs look for?

The top searched job categories for Ai Text To Speech jobs are:

Infographic showing various Ai Text To Speech job openings in the United States as of August 2026, with employment types broken down into 50% Full Time, and 50% Part Time. Highlights an 83% In-person, and 17% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Senior AI Engineer - Voice

Mira Mace

San Francisco, CA โ€ข On-site

$123K - $169K/yr

Full-time

Re-posted 24 days ago


Job description

ABOUT US
Mira Mace pairs Medicare beneficiaries with a dedicated healthcare advocate who navigates appointments, insurance, and care coordination on their behalf. Our customers get the support of caring nurses while AI agents handle the tedious backend work - all covered by Medicare.
We've felt the pain ourselves - the endless back-and-forth with insurance, surprise bills, and the lack of clarity when you just need answers. Too many people fall through the cracks, and we're determined to change that.
Today, 24/7 personalized health assistance is only available to the rich or extremely sick. Our vision is for everyone to be able to afford a health assistant who knows your health history deeply, navigates the healthcare system on your behalf, and propels you to become the healthiest version of yourself.
Our founding team brings a mix of strong technical experience from companies like Google, Meta, Dropbox, and Amazon, along with serial startup experience ranging from early bootstrapped ventures to Series D scale-ups. We are backed by Foundation Capital, DefineVC and top Silicon Valley angel investors.
WHAT WE'RE LOOKING FOR
We're looking for an engineer to own our voice AI end to end - the calling pipeline, and the quality of every conversation that runs through it.
Our problem space runs on the phone. Getting a customer set up takes a real conversation, and much of the work we do on their behalf can only happen by calling institutions - which means automated phone menus, long holds, transfers, and getting a specific outcome out of someone who didn't expect the call. Many of our customers are older and expect to be spoken with, not processed. Voice AI has to clear that bar, and clearing it is the product.
You'll own the full real-time stack and the systems that make it measurably better with every release: evaluation for calls, replayable regression suites, and the tooling that tells us why a call went wrong before a customer does. You'll work directly with the founders and with the team that listens to these calls every day.
This role is specifically for someone with real voice experience. If you've shipped a production real-time voice agent, we want to talk.
RESPONSIBILITIES
  • Own the voice pipeline. Build and operate the real-time stack - telephony, streaming speech-to-text, LLM orchestration, text-to-speech, endpointing, barge-in, and the latency budget that keeps a conversation feeling human.
  • Drive call quality relentlessly. Define what a good call is, measure it, and move the number. Build the evaluation harness - replayable audio and transcripts, regression suites, per-turn scoring - so quality changes are provable rather than anecdotal.
  • Handle the real world. Automated phone menus and DTMF, holds and transfers, voicemail, dropped calls and retries, noisy lines, accents, interruptions, and speaking clearly with an older audience.
  • Scale calling volume reliably. Make the system hold up across thousands of calls a day, with clean handoffs to a human when the agent should stop.
  • Close the loop with the people who hear every call. Turn what our team observes on real conversations into prompt, model, and pipeline changes, fast.
  • Ship fast and iterate with real users. Deploy to production, listen to real conversations, and improve from them - in days, not quarters.
  • Lay the foundation for scale. Make architectural decisions in the voice stack that hold up from hundreds of calls a day to hundreds of thousands.
QUALIFICATIONS
  • You've built and shipped real-time voice AI to production - not a demo, but a system that took real calls with real people.
  • Hands-on depth across the voice stack: streaming speech-to-text, text-to-speech, turn detection and endpointing, barge-in, and latency optimization under a real-time budget.
  • Experience running telephony in production (SIP/PSTN, Twilio or similar) and the failure modes that come with it.
  • You've built evaluation for voice or conversational systems.
  • You've debugged AI systems in production and can trace a bad outcome through audio, transcript, and model behavior to the actual cause.
  • You're comfortable with ambiguity. The playbook doesn't exist yet, and you're excited to build it.
NICE TO HAVE
  • Experience with real-time voice agent frameworks.
  • Speech model fine-tuning, custom vocabularies, or domain adaptation for specialized terminology.
  • Experience building agents that navigate automated phone systems on the other end of the line.
  • Experience in healthcare or another regulated industry (HIPAA, PII/PHI handling, call recording and consent).
WHO YOU ARE
Beyond technical skills, we're looking for someone who embodies the attributes that make great engineers at an early-stage company:
  • Proactive. You move quickly and take a forceful stand without being abrasive. You act without being told what to do and bring new ideas to the company.
  • Analytically sharp. You structure and process qualitative or quantitative data and draw penetrating insights. You learn quickly and absorb new information with ease.
  • High standards with attention to detail. You expect nothing short of the best from yourself and your team. You don't let important details slip through the cracks or derail a project.
  • Passionate and open. You exhibit enthusiasm and a can-do attitude over your work. You solicit feedback often and react calmly to criticism or negative feedback.
OUR CULTURE
Everything we do is guided by a set of leadership principles that define how we operate:
  • Patient First. We start with the patient and work backwards. Every decision - what we build, who we partner with, how we operate - is filtered through one question: does this make the patient's life better?
  • Sense of Urgency. Every day a patient goes unnavigated is a day someone needing help couldn't get the support they needed. We move fast, make decisions with conviction, and carry urgency toward the long-term vision: an AI nurse concierge in every patient's corner.
  • Ownership. We see things through. We don't ship and walk away - we own outcomes, not just tasks. We act on behalf of the entire company, beyond just our own area of responsibility. We never say "that's not my job."
  • Insist on the Highest Standards. We hold ourselves and our teams to nothing short of the best - in clinical quality, in operational execution, in how we show up for patients and partners. We raise the bar continuously.
  • Question Everything, Unapologetically. We reason from first principles, not precedent. We challenge requirements regardless of who set them, dig until we reach the root of the problem, and resist the pull of "that's how it's always been done."
WHY JOIN US
  • Mission with massive impact. Every call your system handles is someone who got the help they needed instead of falling through the cracks. We're building one of the largest AI-first companies in healthcare.
  • Voice is the product. This isn't a feature team. Phone conversations are how the work gets done in this business, and the quality of them is the company's core capability. You own it.
  • Learn fast, build fast. We believe in experimentation, measurement, and steady improvement. You'll ship in days, not quarters.
  • Grow with us. You'll be part of the team that takes Mira Mace from early product to scale. The decisions you make now will define the company's technical DNA.
  • Meaningful early equity. Competitive compensation and real ownership in what we're building.