1

Text To Speech Jobs (NOW HIRING)

Showing results 41-60

Text To Speech information

What is a text to speech job?

A Text To Speech (TTS) job typically involves converting written text into spoken audio using specialized software or AI technology. Professionals in this field may work on developing, fine-tuning, or implementing TTS systems for various applications, such as virtual assistants, accessibility tools, or audiobooks. The role can also include tasks like voice data collection, script editing, and quality assurance of generated speech. TTS jobs are important for making digital content more accessible to people with visual impairments or reading difficulties. The field combines elements of linguistics, software engineering, and artificial intelligence.

What are the key skills and qualifications needed to thrive as a text to speech engineer?

To thrive as a Text to Speech Engineer, you need a strong background in computer science, linguistics, and digital signal processing, often supported by a relevant degree. Experience with machine learning frameworks, speech synthesis toolkits (like Tacotron or WaveNet), and programming languages such as Python or C++ is typically required. Creativity, analytical thinking, and cross-functional communication skills help you collaborate with diverse teams and innovate in voice technology. These skills ensure the development of accurate, natural-sounding speech systems that meet user and client needs.

What are some common challenges faced by professionals working in text to speech development roles?

Professionals in Text to Speech development often encounter challenges such as fine-tuning synthetic voices to sound natural and expressive, handling diverse accents or languages, and optimizing algorithms for various platforms. Collaboration with linguists, UX designers, and software engineers is frequent, as ensuring accessibility and seamless integration across applications is a top priority. Staying updated on advances in AI and deep learning is essential, as the field evolves rapidly and demands continuous improvement in both technical and creative aspects.

What is the difference between Text To Speech vs Voice Actor?

AspectText To SpeechVoice Actor
Required CredentialsNone or basic audio editing skillsVoice training, acting skills, often professional demos
Work EnvironmentSoftware, digital platforms, remoteRecording studios, on-location, remote
Industry UsageAutomation, AI, tech companiesMedia, entertainment, advertising
Search & Comparison IntentAutomated voice solutions, TTS technologyVoice acting, narration, character voices

Text To Speech involves using software to convert written text into spoken words, primarily for automation and digital applications. Voice Actors, on the other hand, provide human voice recordings for media, entertainment, and advertising. While TTS is tech-driven and often used in AI and accessibility tools, Voice Actors bring emotional nuance and personality to their performances. Both roles are essential in their respective industries, but they differ significantly in skills, environment, and purpose.

More about Text To Speech jobs

What cities are hiring for Text To Speech jobs?

Cities with the most Text To Speech job openings:

What states have the most Text To Speech jobs?

States with the most job openings for Text To Speech jobs include:

Infographic showing various Text To Speech job openings in the United States as of August 2026, with employment types broken down into 62% Full Time, 16% Part Time, 6% Temporary, and 16% Contract. Highlights an 84% In-person, and 16% Remote job distribution.

Southern American English Voice Actor (AI Speech & Voice Modeling)

Remote

$50/hr

Other

Re-posted 20 days ago


Job description

This role is for one of our clients

Compensation: $50 per hour

Join an innovative AI research initiative focused on developing next-generation text-to-speech (TTS) systems capable of producing natural, expressive, and human-like voices. We are seeking experienced voice professionals with authentic Southern American English accents to contribute high-quality voice recordings used to train and evaluate advanced speech generation models.

This opportunity is ideal for voice actors, narrators, broadcasters, and other audio professionals who can consistently deliver clear, engaging, and emotionally expressive performances across a variety of scripts.

Important: Voice recordings will be used exclusively to support an internal customer experience (CX) AI agent. Recordings will not be sold, licensed, or repurposed for unrelated products or datasets.

Requirements

Key Responsibilities Record Professional Voice Samples
  • Record high-quality audio across a variety of script types, including conversational, narrative, instructional, and customer interaction scenarios.
  • Deliver natural, expressive speech with excellent control of tone, pacing, pronunciation, and clarity.
  • Maintain consistent voice quality, accent, and delivery throughout multiple recording sessions.
Follow Recording Standards
  • Adhere to detailed recording guidelines regarding recording environment, microphone setup, audio quality, and file formatting.
  • Produce multiple takes with different emotional styles, emphasis, or delivery when requested.
  • Ensure recordings are clean, free from background noise, and meet project quality standards.
Required Qualifications
  • Native Southern American English speaker currently residing in the United States.
  • Authentic Southern U.S. accent (including regions such as Texas, Georgia, Tennessee, North Carolina, South Carolina, Louisiana, Alabama, Mississippi, or similar).
  • Professional experience in voice acting, narration, broadcasting, dubbing, podcasting, or related voice-based work.
  • Access to a professional or near-professional home recording setup, including:
    • High-quality microphone
    • Quiet recording environment
    • Pop filter or equivalent equipment
  • Strong vocal control with excellent diction, pronunciation, intonation, and emotional range.
  • Ability to accurately follow scripts while maintaining a natural and engaging delivery.
  • Availability to contribute 5–10 hours per week during the project period (approximately 1–2 weeks).
Preferred Qualifications
  • Experience recording content for text-to-speech systems, AI speech datasets, audiobooks, IVR platforms, or voice assistants.
  • Familiarity with audio editing software such as Audacity, Adobe Audition, Reaper, or similar tools.
  • Ability to perform multiple vocal styles, including conversational, professional, energetic, calm, and instructional tones.
Important Information
  • Voice recordings may be used to create a voice model for an internal customer experience (CX) AI application.
  • Applicants should only apply if they are comfortable with voice cloning for this specific internal use case.
  • Voice data will not be sold, licensed, or reused for unrelated commercial products or datasets.
Why Join
  • Contribute to the development of next-generation AI speech technologies.
  • Apply your professional voice talent to innovative AI research.
  • Work remotely with a flexible schedule.
  • Participate in a high-impact project helping improve natural, human-like AI voice systems.
Equal Opportunity Statement

We are committed to providing equal opportunities to all qualified applicants without regard to legally protected characteristics. Reasonable accommodations are available upon request.