1

Automatic Speech Recognition Jobs (NOW HIRING)

Speech Recognition (Automatic Speech Recognition - ASR) * Speech Synthesis (Text-to-Speech - TTS) * Voice Activity Detection (VAD) * Audio streaming and low-latency communication * Professional ...

Speech Recognition (Automatic Speech Recognition - ASR) * Speech Synthesis (Text-to-Speech - TTS) * Voice Activity Detection (VAD) * Audio streaming and low-latency communication * Professional ...

Parlance delivers speech recognition as a managed service. That means we blend intelligent speech technologies, including Automatic Speech Recognition and Natural Language Understanding to transform ...

Deep experience with automatic speech recognition (ASR) * Experience training or fine-tuning Whisper, Conformer, wav2vec, or similar speech architectures * Experience with large-scale audio datasets ...

Deep experience with automatic speech recognition (ASR) * Experience training or fine-tuning Whisper, Conformer, wav2vec, or similar speech architectures * Experience with large-scale audio datasets ...

next page

Showing results 1-20

Automatic Speech Recognition information

See salary details

$15

$43

$69

How much do automatic speech recognition jobs pay per hour?

As of Sep 13, 2026, the average hourly pay for automatic speech recognition in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is automatic speech recognition?

Automatic Speech Recognition (ASR) is a technology that enables computers to interpret and transcribe spoken language into text. It uses algorithms and machine learning models to process audio signals, recognize words, and convert them into written form. ASR is widely used in applications such as virtual assistants, transcription services, and voice-activated devices. The technology has advanced significantly in recent years, making voice interactions with computers more accurate and accessible.

What are some common challenges faced by professionals working in automatic speech recognition roles?

Professionals in Automatic Speech Recognition often encounter challenges such as handling diverse accents, background noise, and varying speech patterns, which can affect recognition accuracy. Additionally, ensuring that ASR systems perform well across different languages and specialized vocabularies requires continuous data collection and model training. Collaboration with linguists, software engineers, and product managers is typical, as ASR specialists work to integrate models into real-world applications and improve user experience. Staying updated with the latest research and advancements in machine learning and natural language processing is also essential for career growth in this field.

What are the key skills and qualifications needed to thrive as an automatic speech recognition engineer, and why are they important?

To excel as an Automatic Speech Recognition (ASR) Engineer, you need a solid background in computer science, machine learning, and signal processing, often supported by a relevant degree. Proficiency with Python, deep learning frameworks (such as TensorFlow or PyTorch), and experience with ASR toolkits like Kaldi or ESPnet are typically required. Strong analytical thinking, problem-solving abilities, and effective communication skills distinguish top professionals in this field. These competencies are crucial for developing accurate, efficient ASR systems that drive advancements in voice-enabled technologies.

What is the difference between Automatic Speech Recognition vs Speech Data Annotator?

AspectAutomatic Speech RecognitionSpeech Data Annotator
Required CredentialsComputer science, linguistics, or engineering degree; programming skillsBasic understanding of speech data; attention to detail
Work EnvironmentTech companies, R&D labs, software development teamsData labeling centers, remote or on-site annotation teams
Industry UsageDeveloping speech recognition systems, AI applicationsPreparing datasets for training speech models
Common Search/ComparisonYesNo

Automatic Speech Recognition involves developing systems that convert spoken language into text using advanced algorithms and machine learning. Speech Data Annotators focus on labeling and preparing speech datasets to train these systems. While ASR engineers design the models, annotators provide the essential data for training. Both roles are vital in the speech technology industry, but they differ in skills, responsibilities, and work environments.

More about Automatic Speech Recognition jobs

What cities are hiring for Automatic Speech Recognition jobs?

Cities with the most Automatic Speech Recognition job openings:

What states have the most Automatic Speech Recognition jobs?

States with the most job openings for Automatic Speech Recognition jobs include:

What other helpful pages are available for Automatic Speech Recognition?

Other pages related to Automatic Speech Recognition:

Infographic showing various Automatic Speech Recognition job openings in the United States as of September 2026, with employment types broken down into 1% As Needed, 83% Full Time, 14% Part Time, and 2% Contract. Highlights an 58% Physical, 1% Hybrid, and 41% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Speech Recognition Engineer

Manhattan, NY • On-site

Other

Re-posted 8 days ago


Job description

Job Title: Speech Recognition Engineer
Job Summary
We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled applications. The ideal candidate will have expertise in speech processing, deep learning, natural language processing (NLP), and machine learning. This role involves building, training, fine-tuning, and deploying speech recognition models that deliver high accuracy, low latency, and robust performance across diverse languages, accents, and acoustic environments.
Key Responsibilities
  • Design, develop, and optimize Automatic Speech Recognition (ASR) models for production applications.
  • Build end-to-end speech processing pipelines, including audio preprocessing, feature extraction, decoding, and post-processing.
  • Train, fine-tune, and evaluate speech recognition models using large-scale speech datasets.
  • Improve recognition accuracy for multilingual, domain-specific, and noisy audio environments.
  • Develop real-time and batch speech recognition solutions.
  • Optimize models for latency, throughput, memory efficiency, and inference performance.
  • Integrate ASR models into voice assistants, conversational AI systems, call center platforms, and enterprise applications.
  • Develop data pipelines for speech data collection, annotation, augmentation, and quality validation.
  • Evaluate model performance using industry-standard speech recognition metrics.
  • Collaborate with NLP Engineers, Machine Learning Engineers, AI Engineers, Data Scientists, and Product teams.
  • Deploy speech recognition models using MLOps and cloud-native deployment practices.
  • Monitor production performance and continuously improve model quality.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, Speech Technology, or a related field.
  • 3+ years of experience in speech recognition, speech processing, machine learning, or AI engineering.
  • Strong programming skills in Python.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Solid understanding of digital signal processing (DSP) fundamentals.
  • Experience with speech processing libraries such as SpeechBrain, ESPnet, Hugging Face Transformers, torchaudio, librosa, or Kaldi.
  • Experience training and fine-tuning deep learning models.
  • Familiarity with Linux development environments, Git, and containerization using Docker.
  • Understanding of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
Preferred Qualifications
  • Experience with modern ASR architectures such as Whisper, Conformer, wav2vec 2.0, DeepSpeech, or RNN-Transducer (RNN-T).
  • Experience deploying speech recognition models using ONNX Runtime, TensorRT, NVIDIA Triton Inference Server, or TorchServe.
  • Knowledge of multilingual and low-resource language speech recognition.
  • Experience with streaming speech recognition and real-time inference.
  • Familiarity with speech enhancement, voice activity detection (VAD), speaker diarization, and keyword spotting.
  • Experience with MLOps tools such as MLflow, Kubeflow, or cloud AI platforms.
  • Knowledge of Large Language Models (LLMs) for speech understanding and conversational AI.
Technical Skills
  • Python
  • PyTorch
  • TensorFlow
  • Hugging Face Transformers
  • SpeechBrain
  • ESPnet
  • Kaldi
  • torchaudio
  • librosa
  • Whisper
  • wav2vec 2.0
  • Conformer
  • RNN-T
  • ONNX Runtime
  • TensorRT
  • NVIDIA Triton Inference Server
  • TorchServe
  • Docker
  • Git
  • Linux
  • AWS / Azure / Google Cloud Platform
Soft Skills
  • Strong analytical and problem-solving skills
  • Excellent communication and collaboration
  • Attention to detail
  • Ability to work with cross-functional teams
  • Continuous learning mindset
  • Strong documentation and experimentation practices
Nice to Have
  • Experience with speech synthesis (Text-to-Speech) or conversational AI platforms
  • Knowledge of multilingual ASR evaluation and benchmarking
  • Experience with edge AI deployment for speech applications
  • Familiarity with model compression, quantization, and inference optimization
  • Publications or contributions in speech AI, ASR, or related open-source projects
Key Performance Indicators (KPIs)
  • Word Error Rate (WER) and Character Error Rate (CER)
  • Model inference latency and throughput
  • Speech recognition accuracy across languages and accents
  • Production model availability and reliability
  • Improvement in recognition quality over baseline models
  • Successful deployment and adoption of ASR features
  • Reduction in production defects and model regressions

Location
Hybrid / Remote / On-site (as applicable)
Employment Type
Full-time