Speech Recognition Engineer
Manhattan, NY · On-site +1
Job Title: Speech Recognition Engineer Job Summary We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled ...

Manhattan, NY · On-site +1
Job Title: Speech Recognition Engineer Job Summary We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled ...
Manhattan, NY · On-site +1
Job Title: Speech Recognition Engineer Job Summary We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled ...
San Jose, CA · On-site
As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...
San Jose, CA · On-site
As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...
San Jose, CA · On-site
As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...
San Jose, CA · On-site
As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...
Seattle, WA · On-site
$332K/yr
In this role, you will develop state-of-the-art automatic speech recognition system and ship it to ... As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI ...
Seattle, WA · On-site
$332K/yr
In this role, you will develop state-of-the-art automatic speech recognition system and ship it to ... As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI ...
San Jose, CA · On-site
$151.80 - $332.20/hr
In this role, you will develop state-of-the-art automatic speech recognition system and ship it to ... As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI ...
San Jose, CA · On-site
$151.80 - $332.20/hr
In this role, you will develop state-of-the-art automatic speech recognition system and ship it to ... As an AI Inference Engineer, you will develop novel speech model inference solutions on modern AI ...
Summary Requirement Voice Recognition(VR) implemented across Chrome, Edge, Safari, Firefox, Brave. (required) Customize Web Speech API and other speech frameworks for product needs Other Speech ...
Summary Requirement Voice Recognition(VR) implemented across Chrome, Edge, Safari, Firefox, Brave. (required) Customize Web Speech API and other speech frameworks for product needs Other Speech ...
... and engineering that makes Hippocratic AI's conversational platform not just intelligent, but ... This role exists because accurate speech recognition in clinical contexts is a frontier problem-no ...
... and engineering that makes Hippocratic AI's conversational platform not just intelligent, but ... This role exists because accurate speech recognition in clinical contexts is a frontier problem-no ...
Los Angeles, CA · On-site
$140 - $220/hr
This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... Collaborate with product managers, AI engineers, and software developers to deliver production ...
Los Angeles, CA · On-site
$140 - $220/hr
This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... Collaborate with product managers, AI engineers, and software developers to deliver production ...
This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... If you're passionate about conversational AI and enjoy solving challenging engineering problems, we ...
This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... If you're passionate about conversational AI and enjoy solving challenging engineering problems, we ...
Los Angeles, CA · On-site
This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... If you're passionate about conversational AI and enjoy solving challenging engineering problems, we ...
Los Angeles, CA · On-site
This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... If you're passionate about conversational AI and enjoy solving challenging engineering problems, we ...
Mountain View, CA · On-site
$120 - $170/hr
Train, deploy and maintain scalable automatic speech recognition pipelines to power Level AI's ASR ... tier-1 engineering institute. * Experience deploying Deep Learning based models for speech ...
New
Mountain View, CA · On-site
$120 - $170/hr
Train, deploy and maintain scalable automatic speech recognition pipelines to power Level AI's ASR ... tier-1 engineering institute. * Experience deploying Deep Learning based models for speech ...
New
San Francisco, CA · On-site
$180 - $240/hr
You'll work on cutting-edge problems in medical speech recognition, clinical language understanding ... engineering with focus on NLP/speech recognition * Strong expertise in PyTorch or TensorFlow
San Francisco, CA · On-site
$180 - $240/hr
You'll work on cutting-edge problems in medical speech recognition, clinical language understanding ... engineering with focus on NLP/speech recognition * Strong expertise in PyTorch or TensorFlow
Description The Speech Team within the Siri organization drives major speech recognition, synthesis ... programming skills in Python and one of the deep learning toolkits such as PyTorch, JAX, or ...
Description The Speech Team within the Siri organization drives major speech recognition, synthesis ... programming skills in Python and one of the deep learning toolkits such as PyTorch, JAX, or ...
Cupertino, CA · On-site
$263 - $394/hr
Description The Speech organization within Siri drives major speech recognition, synthesis, and ... Expert programming skills in Python and deep learning frameworks such as PyTorch, JAX, or ...
Cupertino, CA · On-site
$263 - $394/hr
Description The Speech organization within Siri drives major speech recognition, synthesis, and ... Expert programming skills in Python and deep learning frameworks such as PyTorch, JAX, or ...
We're looking for exceptionally skilled and creative scientists and engineers eager to get involved ... such as speech recognition, natural language processing, multi-modal, TTS dialogue management ...
We're looking for exceptionally skilled and creative scientists and engineers eager to get involved ... such as speech recognition, natural language processing, multi-modal, TTS dialogue management ...
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech ... Description: As a Solutions Engineer at Parlance , you deliver meaningful results through a ...
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech ... Description: As a Solutions Engineer at Parlance , you deliver meaningful results through a ...
Cupertino, CA · On-site
$262K - $394K/yr
Description The Speech organization within Siri drives major speech recognition, synthesis, and ... D. in Computer Science, Electrical Engineering, Machine Learning, or similar technical field.
Cupertino, CA · On-site
$262K - $394K/yr
Description The Speech organization within Siri drives major speech recognition, synthesis, and ... D. in Computer Science, Electrical Engineering, Machine Learning, or similar technical field.
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech ... Description: As a Solutions Engineer at Parlance , you deliver meaningful results through a ...
Indiana, PA · Remote
Parlance delivers speech recognition as a managed service. That means we blend intelligent speech ... Description: As a Solutions Engineer at Parlance , you deliver meaningful results through a ...
Tensorflow, PyTorch Experience with speaker recognition and related technologies Curious about new ... engineers eager to get involved in hands-on work to improve our speech technologies by applying ...
Tensorflow, PyTorch Experience with speaker recognition and related technologies Curious about new ... engineers eager to get involved in hands-on work to improve our speech technologies by applying ...
Position Overview We are hiring two ML Engineers / Researchers to help build the next generation of ... Develop and train speech recognition models optimized for medical conversations across hundreds of ...
Position Overview We are hiring two ML Engineers / Researchers to help build the next generation of ... Develop and train speech recognition models optimized for medical conversations across hundreds of ...

Full-time
Re-posted 29 days ago
Design, develop, and optimize Automatic Speech Recognition (ASR) models for production applications.
Build end-to-end speech processing pipelines, including audio preprocessing, feature extraction, decoding, and post-processing.
Train, fine-tune, and evaluate speech recognition models using large-scale speech datasets.
We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled applications. The ideal candidate will have expertise in speech processing, deep learning, natural language processing (NLP), and machine learning. This role involves building, training, fine-tuning, and deploying speech recognition models that deliver high accuracy, low latency, and robust performance across diverse languages, accents, and acoustic environments.
Key ResponsibilitiesDesign, develop, and optimize Automatic Speech Recognition (ASR) models for production applications.
Build end-to-end speech processing pipelines, including audio preprocessing, feature extraction, decoding, and post-processing.
Train, fine-tune, and evaluate speech recognition models using large-scale speech datasets.
Improve recognition accuracy for multilingual, domain-specific, and noisy audio environments.
Develop real-time and batch speech recognition solutions.
Optimize models for latency, throughput, memory efficiency, and inference performance.
Integrate ASR models into voice assistants, conversational AI systems, call center platforms, and enterprise applications.
Develop data pipelines for speech data collection, annotation, augmentation, and quality validation.
Evaluate model performance using industry-standard speech recognition metrics.
Collaborate with NLP Engineers, Machine Learning Engineers, AI Engineers, Data Scientists, and Product teams.
Deploy speech recognition models using MLOps and cloud-native deployment practices.
Monitor production performance and continuously improve model quality.
Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, Speech Technology, or a related field.
3+ years of experience in speech recognition, speech processing, machine learning, or AI engineering.
Strong programming skills in Python.
Experience with deep learning frameworks such as PyTorch or TensorFlow.
Solid understanding of digital signal processing (DSP) fundamentals.
Experience with speech processing libraries such as SpeechBrain, ESPnet, Hugging Face Transformers, torchaudio, librosa, or Kaldi.
Experience training and fine-tuning deep learning models.
Familiarity with Linux development environments, Git, and containerization using Docker.
Understanding of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
Experience with modern ASR architectures such as Whisper, Conformer, wav2vec 2.0, DeepSpeech, or RNN-Transducer (RNN-T).
Experience deploying speech recognition models using ONNX Runtime, TensorRT, NVIDIA Triton Inference Server, or TorchServe.
Knowledge of multilingual and low-resource language speech recognition.
Experience with streaming speech recognition and real-time inference.
Familiarity with speech enhancement, voice activity detection (VAD), speaker diarization, and keyword spotting.
Experience with MLOps tools such as MLflow, Kubeflow, or cloud AI platforms.
Knowledge of Large Language Models (LLMs) for speech understanding and conversational AI.
Python
PyTorch
TensorFlow
Hugging Face Transformers
SpeechBrain
ESPnet
Kaldi
torchaudio
librosa
Whisper
wav2vec 2.0
Conformer
RNN-T
ONNX Runtime
TensorRT
NVIDIA Triton Inference Server
TorchServe
Docker
Git
Linux
AWS / Azure / Google Cloud Platform
Strong analytical and problem-solving skills
Excellent communication and collaboration
Attention to detail
Ability to work with cross-functional teams
Continuous learning mindset
Strong documentation and experimentation practices
Experience with speech synthesis (Text-to-Speech) or conversational AI platforms
Knowledge of multilingual ASR evaluation and benchmarking
Experience with edge AI deployment for speech applications
Familiarity with model compression, quantization, and inference optimization
Publications or contributions in speech AI, ASR, or related open-source projects
Word Error Rate (WER) and Character Error Rate (CER)
Model inference latency and throughput
Speech recognition accuracy across languages and accents
Production model availability and reliability
Improvement in recognition quality over baseline models
Successful deployment and adoption of ASR features
Reduction in production defects and model regressions
Hybrid / Remote / On-site (as applicable)
Employment TypeFull-time