1

Speech Processing Jobs (NOW HIRING)

You will design and implement secure, robust, and scalable services for speech processing; efficient, distributed compute orchestration; optimized scheduling, and more. Your skill at building highly ...

Your work will span multiple use cases and modalities, from reasoning and code generation to text, image, and speech processing. The impact of this role is far-reaching, as you will collaborate cross ...

DSP Engineer - Consumer

Framingham, MA · On-site

$147K - $171K/yr

Demonstrated experience in audio and speech processing such as AEC, noise reduction, speech enhancement, mic array processing or 3D audio. * Exposure to acoustics and acoustic measurements.

DSP Engineer - Consumer

Framingham, MA · On-site

$147K - $171K/yr

Demonstrated experience in audio and speech processing such as AEC, noise reduction, speech enhancement, mic array processing or 3D audio. * Exposure to acoustics and acoustic measurements.

DSP Engineer

Framingham, MA · On-site

$148K - $172K/yr

Demonstrated experience in audio and speech processing such as AEC, noise reduction, speech enhancement, mic array processing or 3D audio.* Exposure to acoustics and acoustic measurements.

Showing results 21-40

Speech Processing information

See salary details

$19

$41

$57

How much do speech processing jobs pay per hour?

As of Sep 13, 2026, the average hourly pay for speech processing in the United States is $41.32, according to ZipRecruiter salary data. Most workers in this role earn between $35.10 and $45.91 per hour, depending on experience, location, and employer.

What is speech processing?

Speech processing is a field within computer science and electrical engineering focused on analyzing, interpreting, and manipulating human speech signals. It includes tasks such as speech recognition, speaker identification, speech synthesis, and speech enhancement. These technologies enable computers and devices to understand spoken language, convert speech to text, or generate natural-sounding synthetic voices. Speech processing is used in virtual assistants, voice-controlled devices, and accessibility tools.

What are the key skills and qualifications needed to thrive as a speech processing engineer, and why are they important?

To excel as a Speech Processing Engineer, you need strong programming skills, a background in digital signal processing, and a degree in computer science, electrical engineering, or a related field. Experience with tools like Python, MATLAB, and libraries such as Kaldi or TensorFlow, as well as knowledge of machine learning frameworks, is typically required. Strong problem-solving, analytical thinking, and clear communication skills help in designing effective solutions and collaborating with cross-functional teams. These skills ensure the development of accurate, efficient speech recognition and synthesis systems that meet user and business needs.

What are some common challenges faced by professionals in speech processing roles, and how can they be addressed?

Professionals in Speech Processing often encounter challenges related to handling diverse accents, background noise, and varying speech patterns, which can impact the accuracy of speech recognition systems. To address these issues, teams frequently collaborate with linguists and data scientists to refine algorithms, utilize large and diverse datasets, and implement advanced noise reduction techniques. Staying updated with the latest research and regularly evaluating system performance are also essential practices to ensure robust and adaptable speech processing solutions.

What is the difference between Speech Processing vs Speech Recognition?

AspectSpeech ProcessingSpeech Recognition
DefinitionBroad field involving analysis, modification, and synthesis of speech signalsSubfield focused on converting spoken language into text
Skills & CertificationsSignal processing, audio engineering, programming; certifications like DSP or audio engineeringMachine learning, NLP, programming; certifications in AI or speech technology
Work EnvironmentResearch labs, tech companies, audio hardware firmsSoftware development, AI companies, voice assistant firms
Industry UsageTelecommunications, audio processing, speech synthesisVirtual assistants, transcription services, voice command systems

Speech Processing is a broad field encompassing various aspects of speech signal analysis and synthesis, while Speech Recognition specifically focuses on converting spoken words into written text. Both roles often require similar technical skills and certifications, but their applications differ across industries and job functions.

More about Speech Processing jobs

What cities are hiring for Speech Processing jobs?

Cities with the most Speech Processing job openings:

What states have the most Speech Processing jobs?

States with the most job openings for Speech Processing jobs include:

What other helpful pages are available for Speech Processing?

Other pages related to Speech Processing:

Infographic showing various Speech Processing job openings in the United States as of September 2026, with employment types broken down into 1% As Needed, 79% Full Time, 16% Part Time, 1% Temporary, 2% Contract, and 1% Nights. Highlights an 91% Physical, 2% Hybrid, and 7% Remote job distribution, with an average salary of $85,951 per year, or $41.3 per hour.

Media Software Engineer, Speech (Senior-Staff Levels)

Sunnyvale, CA • On-site

$180K - $270K/yr

Other

Medical, Dental, Vision, Retirement, PTO

Posted 22 days ago


Key responsibilities

  • Improve the real-time speech and media systems powering live AI conversations.

  • Reduce latency and optimize responsiveness across audio streaming and speech pipelines.

  • Develop new voice and video capabilities that enable more immersive interactions between users and AI bots.


Job description

About Cantina:

Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.

If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!

About the Role:

The Media Team at Cantina is building the real-time infrastructure powering live conversations between people and AI characters. Our goal is simple to express, but challenging to make real: enabling fast, natural, and truly conversational interaction with diverse and creative characters.

We’re looking for a Software Engineer to help improve the speech, audio, and media systems at the heart of the Cantina experience.

This team’s responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based infrastructure supporting real-time voice and video interactions across iOS, Android, and web.

If you’re excited by high-performance C++, real-time systems, speech technologies, and building the future of conversational AI, we’d love to talk.

What You’ll Do:

  • Improve the real-time speech and media systems powering live AI conversations.

  • Reduce latency and optimize responsiveness across audio streaming and speech pipelines.

  • Build tools to enable the creation of AI models to support speech processing.

  • Develop new voice and video capabilities that enable more immersive interactions between users and AI bots.

  • Improve and extend our custom WebRTC infrastructure across iOS, Android, and web.

What You’ll Bring:

Minimum qualifications:

  • BS or MS in Computer Science, Computer Engineering, or a related field; or equivalent experience.

  • 3+ years of experience working as a software engineer.

  • Excellent communications skills.

  • Demonstrated ability to work independently to drive projects from requirements to completion.

  • Experience with C or C++ in a professional context.

  • Grounding in computer science fundamentals, including memory management, high-performance data structures, and concurrent / multithreaded systems.

  • Exposure to system programming concepts, including network protocol design, asynchronous I/O, and distributed system architectures..

  • Object-oriented development and design skills.

  • Interest in solving subtle and challenging engineering problems.

Preferred qualifications:

  • Previous experience with WebRTC, streaming protocols, or other media-adjacent technologies.

  • Familiarity with media processing techniques..

  • Experience creating backend server infrastructure.

  • Experience developing software for iOS or Android.

  • Familiarity with building services using Node.js or Go.

  • Familiarity with artificial intelligence and machine learning techniques, particularly in relation to speech recognition and synthesis.

Compensation:

The anticipated annual base salary range for this role is between $180,000-$270,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.

Benefits:

  • Competitive salary and generous company equity

  • Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina

  • 42 days of paid time off, including:

    • 15 PTO days
    • 10 sick days
    • 15 company holidays
    • 2 floating holidays
  • Generous parental leave & fertility support

  • 401(k) retirement savings plan

  • Lifestyle spending account – $500/month to use however you’d like

  • Complimentary lunch and snacks for in-office employees

  • One Medical membership, and more!

#J-18808-Ljbffr