1

Speech Recognition Engineer Jobs (NOW HIRING)

As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...

As a Speech Recognition Engineer , you will be responsible for consumer product design for our Advanced Products and Technologies. In this role, you will be involved in modeling techniques to advance ...

Parlance delivers speech recognition as a managed service. That means we blend intelligent speech ... Description: As a Solutions Engineer at Parlance , you deliver meaningful results through a ...

Parlance delivers speech recognition as a managed service. That means we blend intelligent speech ... Description: As a Solutions Engineer at Parlance , you deliver meaningful results through a ...

next page

Showing results 1-20

Speech Recognition Engineer information

What is a speech recognition engineer?

A Speech Recognition Engineer develops and optimizes systems that convert spoken language into text using machine learning and signal processing techniques. They work with large datasets to train and fine-tune speech models, improve accuracy, and handle various accents, noises, and languages. Their role often involves programming in Python, using frameworks like TensorFlow or Kaldi, and collaborating with linguists and software engineers to enhance speech-enabled applications.

What does a speech recognition engineer do?

As a Speech Recognition Engineer, you’ll typically work on developing, optimizing, and maintaining ASR models and related tools, often as part of a multidisciplinary team. Your daily responsibilities may include analyzing large audio datasets, researching new model architectures, training and evaluating speech models, and troubleshooting real-world system issues. You’ll also collaborate closely with linguists, software developers, and product managers to integrate speech solutions into larger products or platforms. This role offers opportunities to stay on the cutting edge of AI technology and contribute to impactful, user-facing advancements.

What are the key skills and qualifications needed to thrive as a speech recognition engineer?

To thrive as a Speech Recognition Engineer, a strong background in computer science, signal processing, and machine learning—often backed by a relevant degree—is essential. Familiarity with frameworks such as TensorFlow or PyTorch, knowledge of ASR (Automatic Speech Recognition) systems, and experience with programming languages like Python and C++ are highly valued. Excellent problem-solving skills, teamwork, and effective communication are important soft skills for this position. These competencies enable engineers to build robust, high-quality speech recognition systems that perform well in real-world applications and collaborative work environments.

More about Speech Recognition Engineer jobs
What cities are hiring for Speech Recognition Engineer jobs? Cities with the most Speech Recognition Engineer job openings:
What are the most commonly searched types of Speech Recognition Engineer jobs? The most popular types of Speech Recognition Engineer jobs are:
What states have the most Speech Recognition Engineer jobs? States with the most job openings for Speech Recognition Engineer jobs include:
What job categories do people searching Speech Recognition Engineer jobs look for? The top searched job categories for Speech Recognition Engineer jobs are:
Infographic showing various Speech Recognition Engineer job openings in the United States as of August 2026, with employment types broken down into 2% As Needed, 80% Full Time, 16% Part Time, and 2% Contract. Highlights an 66% Physical, 1% Hybrid, and 33% Remote job distribution.

Speech Recognition Engineer

Ova Technologies

Manhattan, NY • On-site

Other

Re-posted 2 days ago


Job description

Job Title: Speech Recognition Engineer
Job Summary
We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled applications. The ideal candidate will have expertise in speech processing, deep learning, natural language processing (NLP), and machine learning. This role involves building, training, fine-tuning, and deploying speech recognition models that deliver high accuracy, low latency, and robust performance across diverse languages, accents, and acoustic environments.
Key Responsibilities
  • Design, develop, and optimize Automatic Speech Recognition (ASR) models for production applications.
  • Build end-to-end speech processing pipelines, including audio preprocessing, feature extraction, decoding, and post-processing.
  • Train, fine-tune, and evaluate speech recognition models using large-scale speech datasets.
  • Improve recognition accuracy for multilingual, domain-specific, and noisy audio environments.
  • Develop real-time and batch speech recognition solutions.
  • Optimize models for latency, throughput, memory efficiency, and inference performance.
  • Integrate ASR models into voice assistants, conversational AI systems, call center platforms, and enterprise applications.
  • Develop data pipelines for speech data collection, annotation, augmentation, and quality validation.
  • Evaluate model performance using industry-standard speech recognition metrics.
  • Collaborate with NLP Engineers, Machine Learning Engineers, AI Engineers, Data Scientists, and Product teams.
  • Deploy speech recognition models using MLOps and cloud-native deployment practices.
  • Monitor production performance and continuously improve model quality.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, Speech Technology, or a related field.
  • 3+ years of experience in speech recognition, speech processing, machine learning, or AI engineering.
  • Strong programming skills in Python.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Solid understanding of digital signal processing (DSP) fundamentals.
  • Experience with speech processing libraries such as SpeechBrain, ESPnet, Hugging Face Transformers, torchaudio, librosa, or Kaldi.
  • Experience training and fine-tuning deep learning models.
  • Familiarity with Linux development environments, Git, and containerization using Docker.
  • Understanding of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
Preferred Qualifications
  • Experience with modern ASR architectures such as Whisper, Conformer, wav2vec 2.0, DeepSpeech, or RNN-Transducer (RNN-T).
  • Experience deploying speech recognition models using ONNX Runtime, TensorRT, NVIDIA Triton Inference Server, or TorchServe.
  • Knowledge of multilingual and low-resource language speech recognition.
  • Experience with streaming speech recognition and real-time inference.
  • Familiarity with speech enhancement, voice activity detection (VAD), speaker diarization, and keyword spotting.
  • Experience with MLOps tools such as MLflow, Kubeflow, or cloud AI platforms.
  • Knowledge of Large Language Models (LLMs) for speech understanding and conversational AI.
Technical Skills
  • Python
  • PyTorch
  • TensorFlow
  • Hugging Face Transformers
  • SpeechBrain
  • ESPnet
  • Kaldi
  • torchaudio
  • librosa
  • Whisper
  • wav2vec 2.0
  • Conformer
  • RNN-T
  • ONNX Runtime
  • TensorRT
  • NVIDIA Triton Inference Server
  • TorchServe
  • Docker
  • Git
  • Linux
  • AWS / Azure / Google Cloud Platform
Soft Skills
  • Strong analytical and problem-solving skills
  • Excellent communication and collaboration
  • Attention to detail
  • Ability to work with cross-functional teams
  • Continuous learning mindset
  • Strong documentation and experimentation practices
Nice to Have
  • Experience with speech synthesis (Text-to-Speech) or conversational AI platforms
  • Knowledge of multilingual ASR evaluation and benchmarking
  • Experience with edge AI deployment for speech applications
  • Familiarity with model compression, quantization, and inference optimization
  • Publications or contributions in speech AI, ASR, or related open-source projects
Key Performance Indicators (KPIs)
  • Word Error Rate (WER) and Character Error Rate (CER)
  • Model inference latency and throughput
  • Speech recognition accuracy across languages and accents
  • Production model availability and reliability
  • Improvement in recognition quality over baseline models
  • Successful deployment and adoption of ASR features
  • Reduction in production defects and model regressions

Location
Hybrid / Remote / On-site (as applicable)
Employment Type
Full-time