1

Voice Recognition Engineer Jobs (NOW HIRING)

Test voice recognition and recitation outputs for accuracy and clarity * Analyze AI responses for ... Work closely with AI/engineering teams to improve models * Validate multilingual voice data and ...

This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... If you're passionate about conversational AI and enjoy solving challenging engineering problems, we ...

This role involves working with state-of-the-art speech recognition, speech synthesis, voice ... If you're passionate about conversational AI and enjoy solving challenging engineering problems, we ...

With revenues over $500 million, the Branham Group has recognized Procom as the 3rd largest ... Voice Engineer Job Details Demonstrated skills on Avaya systems, which includes. Avaya ...

Title: Voice Dial Plan Engineer Location: Greenwood Village, CO or Denver, CO Type: Fulltime ... Researching new technologies, understanding existing processes, and referencing recognized service ...

Voice Engineer

Fort Collins, CO · On-site

$34.62 - $51.93/hr

... engineers on more complex issues. • Represent and implement voice solutions for new office and ... Recognition * Performance bonus: UCHealth offers a 3-Year Incentive Bonus to recognize employee ...

$120K - $140K/yr

What We're Looking For: We're seeking an experienced and inventive Voice Bot Engineer to lead the ... Experience working with speech recognition (ASR) and text-to-speech (TTS) engines. * Comfort with ...

Voice Engineer

Fort Collins, CO · Hybrid

$34.62 - $51.93/hr

... engineers on more complex issues. Represent and implement voice solutions for new office and ... Recognition * Performance bonus: UCHealth offers a 3-Year Incentive Bonus to recognize employee ...

Voice Bot Engineer

$120K - $140K/yr

What We're Looking For: We're seeking an experienced and inventive Voice Bot Engineer to lead the ... Experience working with speech recognition (ASR) and text-to-speech (TTS) engines. * Comfort with ...

About the role We are looking for a Lead Engineer, AI Agent Voice Experience to help build the next ... Strong background in one or more of the following: speech recognition, speech processing, NLP ...

We are seeking a Software Engineer to join a Global Automotive company in Raymond, OH support the ... Experience with infotainment systems (Bluetooth, Apple CarPlay, Android Auto, Voice Recognition ...

... Developer or other Genesys framework Certifications would be a plus). Qualifications * Minimum of ... Voice Recognition and Design: Nuance Dialog Modules, GRXML, VUI-design o Databases: MS SQL, Oracle ...

next page

Showing results 1-20

Voice Recognition Engineer information

See salary details

$5

$48

$76

How much do voice recognition engineer jobs pay per hour?

As of Sep 11, 2026, the average hourly pay for voice recognition engineer in the United States is $48.17, according to ZipRecruiter salary data. Most workers in this role earn between $39.18 and $60.10 per hour, depending on experience, location, and employer.

What are popular job titles related to Voice Recognition Engineer jobs?

For Voice Recognition Engineer jobs, the most frequently searched job titles are:

Infographic showing various Voice Recognition Engineer job openings in the United States as of June 2026, with employment types broken down into 91% Full Time, 3% Part Time, and 6% Contract. Highlights an 90% Physical, 2% Hybrid, and 8% Remote job distribution, with an average salary of $100,198 per year, or $48.2 per hour.

Speech Recognition Engineer

Manhattan, NY • On-site

Other

Re-posted 6 days ago


Job description

Job Title: Speech Recognition Engineer
Job Summary
We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled applications. The ideal candidate will have expertise in speech processing, deep learning, natural language processing (NLP), and machine learning. This role involves building, training, fine-tuning, and deploying speech recognition models that deliver high accuracy, low latency, and robust performance across diverse languages, accents, and acoustic environments.
Key Responsibilities
  • Design, develop, and optimize Automatic Speech Recognition (ASR) models for production applications.
  • Build end-to-end speech processing pipelines, including audio preprocessing, feature extraction, decoding, and post-processing.
  • Train, fine-tune, and evaluate speech recognition models using large-scale speech datasets.
  • Improve recognition accuracy for multilingual, domain-specific, and noisy audio environments.
  • Develop real-time and batch speech recognition solutions.
  • Optimize models for latency, throughput, memory efficiency, and inference performance.
  • Integrate ASR models into voice assistants, conversational AI systems, call center platforms, and enterprise applications.
  • Develop data pipelines for speech data collection, annotation, augmentation, and quality validation.
  • Evaluate model performance using industry-standard speech recognition metrics.
  • Collaborate with NLP Engineers, Machine Learning Engineers, AI Engineers, Data Scientists, and Product teams.
  • Deploy speech recognition models using MLOps and cloud-native deployment practices.
  • Monitor production performance and continuously improve model quality.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, Speech Technology, or a related field.
  • 3+ years of experience in speech recognition, speech processing, machine learning, or AI engineering.
  • Strong programming skills in Python.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Solid understanding of digital signal processing (DSP) fundamentals.
  • Experience with speech processing libraries such as SpeechBrain, ESPnet, Hugging Face Transformers, torchaudio, librosa, or Kaldi.
  • Experience training and fine-tuning deep learning models.
  • Familiarity with Linux development environments, Git, and containerization using Docker.
  • Understanding of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
Preferred Qualifications
  • Experience with modern ASR architectures such as Whisper, Conformer, wav2vec 2.0, DeepSpeech, or RNN-Transducer (RNN-T).
  • Experience deploying speech recognition models using ONNX Runtime, TensorRT, NVIDIA Triton Inference Server, or TorchServe.
  • Knowledge of multilingual and low-resource language speech recognition.
  • Experience with streaming speech recognition and real-time inference.
  • Familiarity with speech enhancement, voice activity detection (VAD), speaker diarization, and keyword spotting.
  • Experience with MLOps tools such as MLflow, Kubeflow, or cloud AI platforms.
  • Knowledge of Large Language Models (LLMs) for speech understanding and conversational AI.
Technical Skills
  • Python
  • PyTorch
  • TensorFlow
  • Hugging Face Transformers
  • SpeechBrain
  • ESPnet
  • Kaldi
  • torchaudio
  • librosa
  • Whisper
  • wav2vec 2.0
  • Conformer
  • RNN-T
  • ONNX Runtime
  • TensorRT
  • NVIDIA Triton Inference Server
  • TorchServe
  • Docker
  • Git
  • Linux
  • AWS / Azure / Google Cloud Platform
Soft Skills
  • Strong analytical and problem-solving skills
  • Excellent communication and collaboration
  • Attention to detail
  • Ability to work with cross-functional teams
  • Continuous learning mindset
  • Strong documentation and experimentation practices
Nice to Have
  • Experience with speech synthesis (Text-to-Speech) or conversational AI platforms
  • Knowledge of multilingual ASR evaluation and benchmarking
  • Experience with edge AI deployment for speech applications
  • Familiarity with model compression, quantization, and inference optimization
  • Publications or contributions in speech AI, ASR, or related open-source projects
Key Performance Indicators (KPIs)
  • Word Error Rate (WER) and Character Error Rate (CER)
  • Model inference latency and throughput
  • Speech recognition accuracy across languages and accents
  • Production model availability and reliability
  • Improvement in recognition quality over baseline models
  • Successful deployment and adoption of ASR features
  • Reduction in production defects and model regressions

Location
Hybrid / Remote / On-site (as applicable)
Employment Type
Full-time