$180 - $270/hr
This team's responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based ...
$180 - $270/hr
This team's responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based ...
$180 - $270/hr
This team's responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based ...
Sunnyvale, CA · On-site
$180 - $270/hr
This team's responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based ...
Sunnyvale, CA · On-site
$180 - $270/hr
This team's responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based ...
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech technologies, including Automatic Speech Recognition and Natural Language Understanding to transform ...
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech technologies, including Automatic Speech Recognition and Natural Language Understanding to transform ...
Chicago, IL · On-site
$66/hr
Swift UI * Candidates with deep architectural knowledge, strong speech recognition experience, and the ability to reverse engineer complex mobile applications. Overview: We are looking for a Senior ...
Quick apply
Chicago, IL · On-site
$66/hr
Swift UI * Candidates with deep architectural knowledge, strong speech recognition experience, and the ability to reverse engineer complex mobile applications. Overview: We are looking for a Senior ...
Key job responsibilities As a highly experienced science leader in speech and audio processing, you will apply state-of-the-art research in automatic speech recognition, voice synthesis, speech ...
Key job responsibilities As a highly experienced science leader in speech and audio processing, you will apply state-of-the-art research in automatic speech recognition, voice synthesis, speech ...
Key job responsibilities As a highly experienced science leader in speech and audio processing, you will apply state-of-the-art research in automatic speech recognition, voice synthesis, speech ...
Key job responsibilities As a highly experienced science leader in speech and audio processing, you will apply state-of-the-art research in automatic speech recognition, voice synthesis, speech ...
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech technologies, including Automatic Speech Recognition and Natural Language Understanding to transform ...
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech technologies, including Automatic Speech Recognition and Natural Language Understanding to transform ...
Tampa, FL · On-site
Knowledge of speech recognition and synthesis technologies. • Computer Vision: Experience with computer vision tasks such as image classification, object detection, and segmentation. • (Retrieval ...
Tampa, FL · On-site
Knowledge of speech recognition and synthesis technologies. • Computer Vision: Experience with computer vision tasks such as image classification, object detection, and segmentation. • (Retrieval ...
Colorado Springs, CO · On-site
Speech VUI design and documentation. Interactive Voice Response (IVR) application requirements discovery/documentation with experience in Speech Recognition design. Must have excellent skills in ...
Colorado Springs, CO · On-site
Speech VUI design and documentation. Interactive Voice Response (IVR) application requirements discovery/documentation with experience in Speech Recognition design. Must have excellent skills in ...
Cupertino, CA · On-site
$151K - $199K/yr
As a software engineer on the Siri speech team, you will be an important member of a diverse and ... recognition or other machine learning technologies
Cupertino, CA · On-site
$151K - $199K/yr
As a software engineer on the Siri speech team, you will be an important member of a diverse and ... recognition or other machine learning technologies
... Speech Recognition), TTS (Text-to-Speech), and NLP/LLM pipelines. • Create frameworks for conversational flows, prompt engineering, retrieval-augmented generation (RAG), and context management.
... Speech Recognition), TTS (Text-to-Speech), and NLP/LLM pipelines. • Create frameworks for conversational flows, prompt engineering, retrieval-augmented generation (RAG), and context management.
Manhattan, NY · On-site
$174 - $252/hr
Bachelor's degree in Computer Science, Speech Recognition, Computational Linguistics, a related technical field, or equivalent practical experience. * Experience conducting research or development in ...
Manhattan, NY · On-site
$174 - $252/hr
Bachelor's degree in Computer Science, Speech Recognition, Computational Linguistics, a related technical field, or equivalent practical experience. * Experience conducting research or development in ...
Bachelor's degree in Computer Science, Speech Recognition, Computational Linguistics, a related technical field, or equivalent practical experience. * Experience conducting research or development in ...
Bachelor's degree in Computer Science, Speech Recognition, Computational Linguistics, a related technical field, or equivalent practical experience. * Experience conducting research or development in ...
Chicago, IL · On-site
... Speech Recognition, Genesys Voice Portal. • Experience with implementing complex solutions involving Web Services, SOAP, VXML and Voice Object • Provide technical expertise in all aspects of IVR ...
Chicago, IL · On-site
... Speech Recognition, Genesys Voice Portal. • Experience with implementing complex solutions involving Web Services, SOAP, VXML and Voice Object • Provide technical expertise in all aspects of IVR ...
Lenexa, KS · On-site
$17.25 - $21/hr
They may also review and edit medical documents created using speech recognition technology. Transcriptionists interpret medical terminology and abbreviations in preparing patients' medical histories ...
Lenexa, KS · On-site
$17.25 - $21/hr
They may also review and edit medical documents created using speech recognition technology. Transcriptionists interpret medical terminology and abbreviations in preparing patients' medical histories ...
Bellevue, WA · On-site
$123 - $230/hr
Develop general‑purpose, end‑to‑end large speech models covering multilingual automatic speech recognition (ASR), speech translation, speech synthesis, paralinguistic understanding, and general ...
New
Bellevue, WA · On-site
$123 - $230/hr
Develop general‑purpose, end‑to‑end large speech models covering multilingual automatic speech recognition (ASR), speech translation, speech synthesis, paralinguistic understanding, and general ...
New
Cupertino, CA · On-site
$180 - $260/hr
Speech Scientist / Engineer (Interspeech 2022) Cupertino, California, United States Software and ... Experience with speaker recognition and related technologies * Curious about new technologies and ...
New
Cupertino, CA · On-site
$180 - $260/hr
Speech Scientist / Engineer (Interspeech 2022) Cupertino, California, United States Software and ... Experience with speaker recognition and related technologies * Curious about new technologies and ...
New
Key job responsibilities As a highly experienced science leader in speech and audio processing, you will apply state-of-the-art research in automatic speech recognition, voice synthesis, speech ...
Key job responsibilities As a highly experienced science leader in speech and audio processing, you will apply state-of-the-art research in automatic speech recognition, voice synthesis, speech ...
San Jose, CA · On-site
$180K - $450K/yr
Drive research and development to advance speech and audio capabilities in multimodal models, including speech recognition, synthesis, and understanding. * Develop and improve large-scale speech and ...
San Jose, CA · On-site
$180K - $450K/yr
Drive research and development to advance speech and audio capabilities in multimodal models, including speech recognition, synthesis, and understanding. * Develop and improve large-scale speech and ...
$107K - $146K/yr
Responsibilities : • Design, train, fine-tune, and evaluate ML models for speech recognition, generative and reasoning models, and multimodal inference • Adapt open-source and foundation models ...
$107K - $146K/yr
Responsibilities : • Design, train, fine-tune, and evaluate ML models for speech recognition, generative and reasoning models, and multimodal inference • Adapt open-source and foundation models ...
$15.63 - $20.54
0% of jobs
$20.54 - $25.46
2% of jobs
$25.46 - $30.38
6% of jobs
$30.38 - $35.29
12% of jobs
$36.60 is the 25th percentile. Wages below this are outliers.
$35.29 - $40.21
17% of jobs
The median wage is $43.54 / hr.
$40.21 - $45.13
18% of jobs
$45.13 - $50.04
17% of jobs
$50.76 is the 75th percentile. Wages above this are outliers.
$50.04 - $54.96
13% of jobs
$54.96 - $59.88
8% of jobs
$59.88 - $64.79
4% of jobs
$64.79 - $69.71
2% of jobs
$15
$43
$69
A Speech Recognition job involves developing and improving systems that convert spoken language into text. Professionals in this field work with machine learning, natural language processing (NLP), and signal processing to enhance speech-to-text accuracy. They may design models, train algorithms, and fine-tune systems for applications like virtual assistants, transcription services, and accessibility tools. Strong programming skills and knowledge of AI technologies are essential for success in this role.
On a daily basis, speech recognition professionals are typically involved in designing, training, and optimizing speech-to-text models, analyzing audio data, and troubleshooting system errors. They often collaborate with software engineers, data scientists, and product teams to integrate speech technologies into various applications. Responsibilities may also include evaluating model performance using large speech datasets, updating acoustic or language models, and keeping up to date with advancements in machine learning techniques. This blend of technical and collaborative work ensures continuous improvement and innovation in speech recognition products.
To excel in Speech Recognition, you need a strong background in computational linguistics, machine learning, and signal processing, often supported by a degree in computer science, engineering, or a related field. Familiarity with tools such as Python, TensorFlow, Kaldi, and speech corpus databases, alongside relevant certifications in AI or data science, is highly valuable. Strong analytical skills, attention to detail, and the ability to collaborate effectively with cross-functional teams are important soft skills. These competencies are critical for developing, refining, and implementing accurate and efficient speech recognition systems that meet real-world needs.
Cities with the most Speech Recognition job openings:
The most popular types of Speech Recognition jobs are:
States with the most job openings for Speech Recognition jobs include:
The top searched job categories for Speech Recognition jobs are:

$180 - $270/hr
Other
Medical, Dental, Vision, Retirement, PTO
Posted 13 days ago
Improve the real-time speech and media systems powering live AI conversations.
Reduce latency and optimize responsiveness across audio streaming and speech pipelines.
Develop tools to enable the creation of AI models to support speech processing.
About Cantina:
Cantina Labs is a social AI company, developing a suite of advanced real-time models that push the boundaries of expression, personality, and realism. We bring characters to life, transforming how people tell stories, connect, and create. We build and power ecosystems. Cantina, our flagship social AI platform, is just the beginning.
If you're excited about the potential AI has to shape human creativity and social interactions, join us in building the future!
About the Role:
The Media Team at Cantina is building the real-time infrastructure powering live conversations between people and AI characters. Our goal is simple to express, but challenging to make real: enabling fast, natural, and truly conversational interaction with diverse and creative characters.
We’re looking for a Software Engineer to help improve the speech, audio, and media systems at the heart of the Cantina experience.
This team’s responsibilities cover low-level media processing pipelines, integration with internal and external models for speech recognition and synthesis, and globally distributed, WebRTC-based infrastructure supporting real-time voice and video interactions across iOS, Android, and web.
If you’re excited by high-performance C++, real-time systems, speech technologies, and building the future of conversational AI, we’d love to talk.
What You’ll Do:
Improve the real-time speech and media systems powering live AI conversations.
Reduce latency and optimize responsiveness across audio streaming and speech pipelines.
Build tools to enable the creation of AI models to support speech processing.
Develop new voice and video capabilities that enable more immersive interactions between users and AI bots.
Improve and extend our custom WebRTC infrastructure across iOS, Android, and web.
What You’ll Bring:
Minimum qualifications:
BS or MS in Computer Science, Computer Engineering, or a related field; or equivalent experience.
3+ years of experience working as a software engineer.
Excellent communications skills.
Demonstrated ability to work independently to drive projects from requirements to completion.
Experience with C or C++ in a professional context.
Grounding in computer science fundamentals, including memory management, high-performance data structures, and concurrent / multithreaded systems.
Exposure to system programming concepts, including network protocol design, asynchronous I/O, and distributed system architectures..
Object-oriented development and design skills.
Interest in solving subtle and challenging engineering problems.
Preferred qualifications:
Previous experience with WebRTC, streaming protocols, or other media-adjacent technologies.
Familiarity with media processing techniques..
Experience creating backend server infrastructure.
Experience developing software for iOS or Android.
Familiarity with building services using Node.js or Go.
Familiarity with artificial intelligence and machine learning techniques, particularly in relation to speech recognition and synthesis.
Compensation:
The anticipated annual base salary range for this role is between $180,000-$270,000. When determining compensation, a number of factors will be considered, including skills, experience, job scope, location, and competitive compensation market data.
Benefits:
Competitive salary and generous company equity
Medical, dental, and vision insurance – 99.99% of premiums covered by Cantina
42 days of paid time off, including:
Generous parental leave & fertility support
401(k) retirement savings plan
Lifestyle spending account – $500/month to use however you’d like
Complimentary lunch and snacks for in-office employees
One Medical membership, and more!