1

Speech Recognition Jobs (NOW HIRING)

$180 - $270/hr

This team's responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based ...

Parlance delivers speech recognition as a managed service. That means we blend intelligent speech technologies, including Automatic Speech Recognition and Natural Language Understanding to transform ...

Parlance delivers speech recognition as a managed service. That means we blend intelligent speech technologies, including Automatic Speech Recognition and Natural Language Understanding to transform ...

Speech VUI design and documentation. Interactive Voice Response (IVR) application requirements discovery/documentation with experience in Speech Recognition design. Must have excellent skills in ...

Senior Engineer (ML/AI)

$107K - $146K/yr

Responsibilities : • Design, train, fine-tune, and evaluate ML models for speech recognition, generative and reasoning models, and multimodal inference • Adapt open-source and foundation models ...

Showing results 41-60

Speech Recognition information

See salary details

$15

$43

$69

How much do speech recognition jobs pay per hour?

As of Sep 4, 2026, the average hourly pay for speech recognition in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is speech recognition?

A Speech Recognition job involves developing and improving systems that convert spoken language into text. Professionals in this field work with machine learning, natural language processing (NLP), and signal processing to enhance speech-to-text accuracy. They may design models, train algorithms, and fine-tune systems for applications like virtual assistants, transcription services, and accessibility tools. Strong programming skills and knowledge of AI technologies are essential for success in this role.

What does a speech recognition professional do?

On a daily basis, speech recognition professionals are typically involved in designing, training, and optimizing speech-to-text models, analyzing audio data, and troubleshooting system errors. They often collaborate with software engineers, data scientists, and product teams to integrate speech technologies into various applications. Responsibilities may also include evaluating model performance using large speech datasets, updating acoustic or language models, and keeping up to date with advancements in machine learning techniques. This blend of technical and collaborative work ensures continuous improvement and innovation in speech recognition products.

What skills and qualifications are needed for speech recognition?

To excel in Speech Recognition, you need a strong background in computational linguistics, machine learning, and signal processing, often supported by a degree in computer science, engineering, or a related field. Familiarity with tools such as Python, TensorFlow, Kaldi, and speech corpus databases, alongside relevant certifications in AI or data science, is highly valuable. Strong analytical skills, attention to detail, and the ability to collaborate effectively with cross-functional teams are important soft skills. These competencies are critical for developing, refining, and implementing accurate and efficient speech recognition systems that meet real-world needs.

More about Speech Recognition jobs

What cities are hiring for Speech Recognition jobs?

Cities with the most Speech Recognition job openings:

What are the most commonly searched types of Speech Recognition jobs?

The most popular types of Speech Recognition jobs are:

What states have the most Speech Recognition jobs?

States with the most job openings for Speech Recognition jobs include:

Infographic showing various Speech Recognition job openings in the United States as of August 2026, with employment types broken down into 91% Full Time, and 9% Part Time. Highlights an 91% In-person, and 9% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Media Software Engineer, Speech (Senior-Staff Levels)

AI Chopping Block

On-site

$180 - $270/hr

Other

Medical, Dental, Vision, Retirement, PTO

Posted 13 days ago


Key responsibilities

  • Improve the real-time speech and media systems powering live AI conversations.

  • Reduce latency and optimize responsiveness across audio streaming and speech pipelines.

  • Develop tools to enable the creation of AI models to support speech processing.


Job description

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

The Media Team at Cantina is building the real-time infrastructure powering live conversations between people and AI characters. Our goal is simple to express, but challenging to make real: enabling fast, natural, and truly conversational interaction with diverse and creative characters.

We’re looking for a Software Engineer to help improve the speech, audio, and media systems at the heart of the Cantina experience.

This team’s responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based infrastructure supporting real-time voice and video interactions across iOS, Android, and web.

If you’re excited by high-performance C++, real-time systems, speech technologies, and building the future of conversational AI, we’d love to talk.

What You’ll Do:

  • Improve the real-time speech and media systems powering live AI conversations.

  • Reduce latency and optimize responsiveness across audio streaming and speech pipelines.

  • Build tools to enable the creation of AI models to support speech processing.

  • Develop new voice and video capabilities that enable more immersive interactions between users and AI bots.

  • Improve and extend our custom WebRTC infrastructure across iOS, Android, and web.

What You’ll Bring:

Minimum qualifications:

  • BS or MS in Computer Science, Computer Engineering, or a related field; or equivalent experience.

  • 3+ years of experience working as a software engineer.

  • Excellent communications skills.

  • Demonstrated ability to work independently to drive projects from requirements to completion.

  • Experience with C or C++ in a professional context.

  • Grounding in computer science fundamentals, including memory management, high-performance data structures, and concurrent / multithreaded systems.

  • Exposure to system programming concepts, including network protocol design, asynchronous I/O, and distributed system architectures..

  • Object-oriented development and design skills.

  • Interest in solving subtle and challenging engineering problems.

Preferred qualifications:

  • Previous experience with WebRTC, streaming protocols, or other media-adjacent technologies.

  • Familiarity with media processing techniques..

  • Experience creating backend server infrastructure.

  • Experience developing software for iOS or Android.

  • Familiarity with building services using Node.js or Go.

  • Familiarity with artificial intelligence and machine learning techniques, particularly in relation to speech recognition and synthesis.

Compensation:

The anticipated annual base salary range for this role is between $180,000-$270,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days
    • 10 sick days
    • 15 company holidays
    • 2 floating holidays
  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

#J-18808-Ljbffr