1

Ai Text To Speech Jobs (NOW HIRING)

... text to speech, speech-to-speech models, and production optimization. You will partner closely with ... Improve the quality of voice AI systems through error analysis, data curation, metric design ...

Full Stack AI Engineer

Plano, TX ยท On-site

$140 - $190/hr

Build conversational AI experiences, including chatbots (text) and voicebots (speech-to-text, text-to-speech) * Design and maintain APIs (REST, GraphQL), microservices, and event-driven architectures

FDE with Amazon Connect

Saint Louis, MO ยท On-site

$73K - $170K/yr

Speech-to-speech (S2S) model integration for real-time voice AI * Text-to-speech (TTS) and speech-to-text (STT) pipeline design and optimization * Latency optimization for conversational voice ...

... Speech to Text * Sound knowledge of GCP cloud platform. Screening questions: 1. Google CCAI & Conversational AI Expertise Can you walk us through your experience working with Google CCAI?

... Speech to Text * Sound knowledge of GCP cloud platform. Screening questions: 1. Google CCAI & Conversational AI Expertise Can you walk us through your experience working with Google CCAI?

This role is for one of our clients Compensation: $50 per hour Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) technology capable of producing ...

FDE with Amazon Connect

Saint Louis, MO ยท On-site

$73K - $145K/yr

Speech-to-speech (S2S) model integration for real-time voice AI * Text-to-speech (TTS) and speech-to-text (STT) pipeline design and optimization * Latency optimization for conversational voice ...

Showing results 41-60

Ai Text To Speech information

See salary details

$15

$43

$69

How much do ai text to speech jobs pay per hour?

As of Sep 6, 2026, the average hourly pay for ai text to speech in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is AI Text to Speech?

AI Text to Speech (TTS) is a technology that uses artificial intelligence to convert written text into spoken words. This technology leverages deep learning and neural networks to produce natural-sounding speech that can closely mimic human voices. AI TTS is commonly used in applications such as virtual assistants, accessibility tools for the visually impaired, audiobooks, and automated customer service. It supports multiple languages and can be customized to different voices and accents.

What are the key skills and qualifications needed to thrive as an AI Text-to-Speech engineer?

To thrive as an AI Text-to-Speech Engineer, you need a strong background in computer science, machine learning, and digital signal processing, typically supported by a relevant degree. Familiarity with tools and frameworks such as TensorFlow, PyTorch, speech synthesis engines, and possibly certification in AI/ML technologies is important. Creativity, problem-solving, and effective collaboration with multidisciplinary teams are crucial soft skills. These abilities enable the development of high-quality, natural-sounding TTS systems that meet user needs and industry standards.

What are some common challenges faced by AI Text-to-Speech specialists, and how are they addressed in a typical work environment?

AI Text-to-Speech specialists often encounter challenges such as ensuring natural-sounding speech synthesis, handling diverse accents, and optimizing for different languages or dialects. Addressing these requires collaborating closely with linguists, data engineers, and software developers to fine-tune models and improve datasets. Regular peer reviews and iterative testing are standard to maintain quality and address edge cases. The work environment is typically cross-functional, fostering open communication to solve problems efficiently.

What is the difference between Ai Text To Speech vs Voice Actor?

AspectAi Text To SpeechVoice Actor
CredentialsNone required, but technical skills helpfulVoice training, acting skills, often a demo reel
Work EnvironmentRemote, software-basedStudio or on-location recording
Industry UsageTechnology, media, customer serviceEntertainment, advertising, narration
Work NatureAutomated voice generation, programmingLive or pre-recorded voice performances

Ai Text To Speech involves using software to convert text into synthetic speech, requiring technical knowledge. Voice actors perform live or recorded voice work, emphasizing acting skills and emotional expression. While Ai TTS is automated and scalable, voice actors provide personalized, nuanced performances. Both roles are essential in media and technology industries, but they differ significantly in skills and work environment.

More about Ai Text To Speech jobs

What cities are hiring for Ai Text To Speech jobs?

Cities with the most Ai Text To Speech job openings:

What states have the most Ai Text To Speech jobs?

States with the most job openings for Ai Text To Speech jobs include:

Infographic showing various Ai Text To Speech job openings in the United States as of August 2026, with employment types broken down into 76% Full Time, 20% Part Time, and 4% Contract. Highlights an 66% Physical, 4% Hybrid, and 30% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Remote | Spanish (Peru-based) AI Voice Recording Consultant Up to $40/hour

24-Mag Llc

Manhattan, NY โ€ข Remote

$40/hr

Part-time

This job post hasย expired 2 days ago.ย Applications are no longer accepted.


Job description

Specialised Part-Time Consulting Opportunity for Peru-Based Voice ProfessionalsWe are sharing a specialised part-time consulting opportunity for Peru-based voice professionals experienced in Peruvian Spanish voice recording, narration, dubbing, broadcast delivery, customer experience-style speech, text-to-speech recording, and expressive audio performance. This role supports current and upcoming remote consulting opportunities focused on high-quality Spanish voice recordings for speech model development, customer experience voice applications, pronunciation consistency, expressive delivery, and audio quality review. Selected professionals will record clear, natural, and consistent voice samples across varied scripts while following detailed recording, formatting, and delivery guidelines.

Key ResponsibilitiesProfessionals in this role may contribute to: Record high-quality Peruvian Spanish voice samples across conversational, narrative, instructional, and customer-support style scriptsDeliver clear, natural, and expressive speech with strong control over tone, pacing, pronunciation, and emphasisMaintain consistency in voice, accent, diction, and delivery across multiple recording sessionsPerform multiple takes with variation in emotion, emphasis, and style when requiredSupport voice recording work for text-to-speech and customer experience voice applicationsFollow scripts precisely while maintaining natural, human-sounding deliveryAdapt vocal style for calm, professional, energetic, corporate, conversational, or instructional use cases depending on project needsEnsure recordings meet project expectations for clarity, consistency, pronunciation quality, and expressive rangeFollow detailed recording guidelines related to environment, microphone setup, file formatting, and submission standardsRecord in a quiet environment using a professional or near-professional audio setupReview recordings for basic quality issues before submission, including background noise, inconsistent volume, unclear diction, or technical interruptionsComplete independent remote project tasks in a flexible, deadline-aware environmentIdeal ProfileStrong candidates may have: Native Peruvian Spanish fluency with a natural Peruvian accentFemale voice profile suited to the project's required voice specificationsCurrent location in Peru depending on project requirementsProven experience in voice acting, dubbing, narration, podcasting, broadcasting, IVR, audiobooks, or similar voice recording workStrong command of intonation, diction, pacing, emotional range, and script interpretationAbility to follow recording instructions precisely while maintaining natural deliveryReliable availability for approximately 5โ€“10 hours per week during the project periodComfort with synthetic voice or voice cloning use for a defined customer experience AI voice applicationEducational BackgroundNo specific degree is required for this opportunity.

Professional experience in voice acting, narration, broadcast, dubbing, audio production, communications, theatre, media, or related fields is highly relevantTraining in vocal performance, acting, pronunciation, audio production, media production, or speech delivery may be valuableExperience working with scripts, recording tools, audio submissions, or performance direction may support project fitNice to HaveExperience recording for text-to-speech, AI voice datasets, audiobooks, IVR systems, customer support voice systems, or guided narrationFamiliarity with audio editing tools such as Audacity, Adobe Audition, Reaper, or similar softwareAbility to deliver multiple vocal styles, including conversational, corporate, energetic, calm, warm, or instructional deliveryAccess to a quality microphone, quiet recording environment, pop filter, and reliable file submission workflowStrong ability to maintain vocal consistency across repeated sessions and varied script typesWhy This OpportunityApply Peruvian Spanish voice acting and narration skills to structured remote project workContribute to high-quality TTS voice recording and customer experience voice model evaluationUse expressive delivery, pronunciation control, and audio consistency in a focused recording environmentWork on flexible assignments aligned with voice performance and recording strengthsRemote structure with competitive hourly compensationContract DetailsIndependent contractor roleFully remote with flexible schedulingEligible professionals should be based in Peru depending on project requirementsPart-time commitment, with an estimated 5โ€“10 hours per week depending on project availability and recording needsCompetitive rates up to $40 per hour depending on experience, voice fit, recording quality, and project scopeRecordings may be used to support a specific customer experience AI voice application, including synthetic voice generation or voice cloning for that defined use caseRecordings are intended for the defined project use case and should not be sold, licensed, or reused for unrelated products, datasets, or purposesCandidates should apply only if they are comfortable with the voice cloning and synthetic voice use described for this projectWeekly payments via Stripe or WiseProjects may be extended, shortened, or adjusted depending on scope and performanceWork will not involve access to confidential or proprietary information from any employer, client, or institutionAbout the PlatformThis opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.