1

Speech Recognition Jobs (NOW HIRING)

As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...

As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...

Speech Recognition Lead

San Mateo, CA ยท On-site

$20.50 - $25.75/hr

Develop and implement wake word detectors, optimizing a local speech recognition engine, and creating an event detector to listen for key sounds such as glass cracking. Preferred candidate has ...

Speech Recognition Lead

San Mateo, CA ยท On-site

$20.50 - $25.75/hr

Develop and implement wake word detectors, optimizing a local speech recognition engine, and creating an event detector to listen for key sounds such as glass cracking. Preferred candidate has ...

Senior Speech Scientist - REMOTE

Sacramento, CA ยท On-site +1

$97K - $133K/yr

Functioning as subject matter expert for all issues relating to speech recognition performance * Supporting sales with proposals and questions around speech recognition Requirements In addition to ...

next page

Showing results 1-20

Speech Recognition information

See salary details

$15

$43

$69

How much do speech recognition jobs pay per hour?

As of Aug 9, 2026, the average hourly pay for speech recognition in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What does a speech recognition professional do?

On a daily basis, speech recognition professionals are typically involved in designing, training, and optimizing speech-to-text models, analyzing audio data, and troubleshooting system errors. They often collaborate with software engineers, data scientists, and product teams to integrate speech technologies into various applications. Responsibilities may also include evaluating model performance using large speech datasets, updating acoustic or language models, and keeping up to date with advancements in machine learning techniques. This blend of technical and collaborative work ensures continuous improvement and innovation in speech recognition products.

What is speech recognition?

A Speech Recognition job involves developing and improving systems that convert spoken language into text. Professionals in this field work with machine learning, natural language processing (NLP), and signal processing to enhance speech-to-text accuracy. They may design models, train algorithms, and fine-tune systems for applications like virtual assistants, transcription services, and accessibility tools. Strong programming skills and knowledge of AI technologies are essential for success in this role.

What skills and qualifications are needed for speech recognition?

To excel in Speech Recognition, you need a strong background in computational linguistics, machine learning, and signal processing, often supported by a degree in computer science, engineering, or a related field. Familiarity with tools such as Python, TensorFlow, Kaldi, and speech corpus databases, alongside relevant certifications in AI or data science, is highly valuable. Strong analytical skills, attention to detail, and the ability to collaborate effectively with cross-functional teams are important soft skills. These competencies are critical for developing, refining, and implementing accurate and efficient speech recognition systems that meet real-world needs.

More about Speech Recognition jobs
What cities are hiring for Speech Recognition jobs? Cities with the most Speech Recognition job openings:
What are the most commonly searched types of Speech Recognition jobs? The most popular types of Speech Recognition jobs are:
What states have the most Speech Recognition jobs? States with the most job openings for Speech Recognition jobs include:
Infographic showing various Speech Recognition job openings in the United States as of August 2026, with employment types broken down into 50% Full Time, and 50% Part Time. Highlights an 100% In-person job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Speech Recognition Engineer

Ova Technologies

Manhattan, NY โ€ข On-site

Other

Re-posted 2 days ago


Job description

Job Title: Speech Recognition Engineer
Job Summary
We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled applications. The ideal candidate will have expertise in speech processing, deep learning, natural language processing (NLP), and machine learning. This role involves building, training, fine-tuning, and deploying speech recognition models that deliver high accuracy, low latency, and robust performance across diverse languages, accents, and acoustic environments.
Key Responsibilities
  • Design, develop, and optimize Automatic Speech Recognition (ASR) models for production applications.
  • Build end-to-end speech processing pipelines, including audio preprocessing, feature extraction, decoding, and post-processing.
  • Train, fine-tune, and evaluate speech recognition models using large-scale speech datasets.
  • Improve recognition accuracy for multilingual, domain-specific, and noisy audio environments.
  • Develop real-time and batch speech recognition solutions.
  • Optimize models for latency, throughput, memory efficiency, and inference performance.
  • Integrate ASR models into voice assistants, conversational AI systems, call center platforms, and enterprise applications.
  • Develop data pipelines for speech data collection, annotation, augmentation, and quality validation.
  • Evaluate model performance using industry-standard speech recognition metrics.
  • Collaborate with NLP Engineers, Machine Learning Engineers, AI Engineers, Data Scientists, and Product teams.
  • Deploy speech recognition models using MLOps and cloud-native deployment practices.
  • Monitor production performance and continuously improve model quality.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, Speech Technology, or a related field.
  • 3+ years of experience in speech recognition, speech processing, machine learning, or AI engineering.
  • Strong programming skills in Python.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Solid understanding of digital signal processing (DSP) fundamentals.
  • Experience with speech processing libraries such as SpeechBrain, ESPnet, Hugging Face Transformers, torchaudio, librosa, or Kaldi.
  • Experience training and fine-tuning deep learning models.
  • Familiarity with Linux development environments, Git, and containerization using Docker.
  • Understanding of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
Preferred Qualifications
  • Experience with modern ASR architectures such as Whisper, Conformer, wav2vec 2.0, DeepSpeech, or RNN-Transducer (RNN-T).
  • Experience deploying speech recognition models using ONNX Runtime, TensorRT, NVIDIA Triton Inference Server, or TorchServe.
  • Knowledge of multilingual and low-resource language speech recognition.
  • Experience with streaming speech recognition and real-time inference.
  • Familiarity with speech enhancement, voice activity detection (VAD), speaker diarization, and keyword spotting.
  • Experience with MLOps tools such as MLflow, Kubeflow, or cloud AI platforms.
  • Knowledge of Large Language Models (LLMs) for speech understanding and conversational AI.
Technical Skills
  • Python
  • PyTorch
  • TensorFlow
  • Hugging Face Transformers
  • SpeechBrain
  • ESPnet
  • Kaldi
  • torchaudio
  • librosa
  • Whisper
  • wav2vec 2.0
  • Conformer
  • RNN-T
  • ONNX Runtime
  • TensorRT
  • NVIDIA Triton Inference Server
  • TorchServe
  • Docker
  • Git
  • Linux
  • AWS / Azure / Google Cloud Platform
Soft Skills
  • Strong analytical and problem-solving skills
  • Excellent communication and collaboration
  • Attention to detail
  • Ability to work with cross-functional teams
  • Continuous learning mindset
  • Strong documentation and experimentation practices
Nice to Have
  • Experience with speech synthesis (Text-to-Speech) or conversational AI platforms
  • Knowledge of multilingual ASR evaluation and benchmarking
  • Experience with edge AI deployment for speech applications
  • Familiarity with model compression, quantization, and inference optimization
  • Publications or contributions in speech AI, ASR, or related open-source projects
Key Performance Indicators (KPIs)
  • Word Error Rate (WER) and Character Error Rate (CER)
  • Model inference latency and throughput
  • Speech recognition accuracy across languages and accents
  • Production model availability and reliability
  • Improvement in recognition quality over baseline models
  • Successful deployment and adoption of ASR features
  • Reduction in production defects and model regressions

Location
Hybrid / Remote / On-site (as applicable)
Employment Type
Full-time