1

Speech Processing Jobs (NOW HIRING)

Senior Machine Learning Engineer (Remote)

New York, NY ยท On-site +1

$114K - $157K/yr

Experience with modern speech processing frameworks * Good dev ops experience * Advanced DSP experience * MongoDB or SQL experience * Experience with unit, integration, and load testing * Experience ...

The Alexa Edge AI team is looking for a passionate, talented, and innovative senior applied scientist with a strong background in deep learning and speech processing techniques. You will have an ...

You will design and implement secure, robust, and scalable services for speech processing; efficient, distributed compute orchestration; optimized scheduling, and more. Your skill at building highly ...

Architect and lead development of real-time speech processing pipeline: Voice Command detection, VAD (voice activity detection), ASR, TTS, noise suppression, echo cancellation, and keyword spotting.

You will design and implement secure, robust, and scalable services for speech processing; efficient, distributed compute orchestration; optimized scheduling, and more. Your skill at building highly ...

Your work will span multiple use cases and modalities, from reasoning and code generation to text, image, and speech processing. The impact of this role is far-reaching, as you will collaborate cross ...

DSP Engineer - Consumer

Framingham, MA ยท On-site

$147K - $171K/yr

Demonstrated experience in audio and speech processing such as AEC, noise reduction, speech enhancement, mic array processing or 3D audio. * Exposure to acoustics and acoustic measurements.

DSP Engineer - Consumer

Framingham, MA

$147K - $171K/yr

Demonstrated experience in audio and speech processing such as AEC, noise reduction, speech enhancement, mic array processing or 3D audio. * Exposure to acoustics and acoustic measurements.

next page

Showing results 1-20

Speech Processing information

See salary details

$19

$41

$57

How much do speech processing jobs pay per hour?

As of Aug 12, 2026, the average hourly pay for speech processing in the United States is $41.32, according to ZipRecruiter salary data. Most workers in this role earn between $35.10 and $45.91 per hour, depending on experience, location, and employer.

What is the difference between Speech Processing vs Speech Recognition?

AspectSpeech ProcessingSpeech Recognition
DefinitionBroad field involving analysis, modification, and synthesis of speech signalsSubfield focused on converting spoken language into text
Skills & CertificationsSignal processing, audio engineering, programming; certifications like DSP or audio engineeringMachine learning, NLP, programming; certifications in AI or speech technology
Work EnvironmentResearch labs, tech companies, audio hardware firmsSoftware development, AI companies, voice assistant firms
Industry UsageTelecommunications, audio processing, speech synthesisVirtual assistants, transcription services, voice command systems

Speech Processing is a broad field encompassing various aspects of speech signal analysis and synthesis, while Speech Recognition specifically focuses on converting spoken words into written text. Both roles often require similar technical skills and certifications, but their applications differ across industries and job functions.

What are the key skills and qualifications needed to thrive as a speech processing engineer, and why are they important?

To excel as a Speech Processing Engineer, you need strong programming skills, a background in digital signal processing, and a degree in computer science, electrical engineering, or a related field. Experience with tools like Python, MATLAB, and libraries such as Kaldi or TensorFlow, as well as knowledge of machine learning frameworks, is typically required. Strong problem-solving, analytical thinking, and clear communication skills help in designing effective solutions and collaborating with cross-functional teams. These skills ensure the development of accurate, efficient speech recognition and synthesis systems that meet user and business needs.

What is speech processing?

Speech processing is a field within computer science and electrical engineering focused on analyzing, interpreting, and manipulating human speech signals. It includes tasks such as speech recognition, speaker identification, speech synthesis, and speech enhancement. These technologies enable computers and devices to understand spoken language, convert speech to text, or generate natural-sounding synthetic voices. Speech processing is used in virtual assistants, voice-controlled devices, and accessibility tools.

What are some common challenges faced by professionals in speech processing roles, and how can they be addressed?

Professionals in Speech Processing often encounter challenges related to handling diverse accents, background noise, and varying speech patterns, which can impact the accuracy of speech recognition systems. To address these issues, teams frequently collaborate with linguists and data scientists to refine algorithms, utilize large and diverse datasets, and implement advanced noise reduction techniques. Staying updated with the latest research and regularly evaluating system performance are also essential practices to ensure robust and adaptable speech processing solutions.
More about Speech Processing jobs
What cities are hiring for Speech Processing jobs? Cities with the most Speech Processing job openings:
What states have the most Speech Processing jobs? States with the most job openings for Speech Processing jobs include:
Infographic showing various Speech Processing job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 80% Full Time, 13% Part Time, 1% Temporary, 4% Contract, and 1% Nights. Highlights an 93% Physical, 2% Hybrid, and 5% Remote job distribution, with an average salary of $85,951 per year, or $41.3 per hour.

Speech Recognition Engineer

Ova Technologies

Manhattan, NY โ€ข On-site

Other

Re-posted 5 days ago


Job description

Job Title: Speech Recognition Engineer
Job Summary
We are seeking a Speech Recognition Engineer to design, develop, and optimize Automatic Speech Recognition (ASR) systems for voice-enabled applications. The ideal candidate will have expertise in speech processing, deep learning, natural language processing (NLP), and machine learning. This role involves building, training, fine-tuning, and deploying speech recognition models that deliver high accuracy, low latency, and robust performance across diverse languages, accents, and acoustic environments.
Key Responsibilities
  • Design, develop, and optimize Automatic Speech Recognition (ASR) models for production applications.
  • Build end-to-end speech processing pipelines, including audio preprocessing, feature extraction, decoding, and post-processing.
  • Train, fine-tune, and evaluate speech recognition models using large-scale speech datasets.
  • Improve recognition accuracy for multilingual, domain-specific, and noisy audio environments.
  • Develop real-time and batch speech recognition solutions.
  • Optimize models for latency, throughput, memory efficiency, and inference performance.
  • Integrate ASR models into voice assistants, conversational AI systems, call center platforms, and enterprise applications.
  • Develop data pipelines for speech data collection, annotation, augmentation, and quality validation.
  • Evaluate model performance using industry-standard speech recognition metrics.
  • Collaborate with NLP Engineers, Machine Learning Engineers, AI Engineers, Data Scientists, and Product teams.
  • Deploy speech recognition models using MLOps and cloud-native deployment practices.
  • Monitor production performance and continuously improve model quality.
Required Qualifications
  • Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Electrical Engineering, Speech Technology, or a related field.
  • 3+ years of experience in speech recognition, speech processing, machine learning, or AI engineering.
  • Strong programming skills in Python.
  • Experience with deep learning frameworks such as PyTorch or TensorFlow.
  • Solid understanding of digital signal processing (DSP) fundamentals.
  • Experience with speech processing libraries such as SpeechBrain, ESPnet, Hugging Face Transformers, torchaudio, librosa, or Kaldi.
  • Experience training and fine-tuning deep learning models.
  • Familiarity with Linux development environments, Git, and containerization using Docker.
  • Understanding of cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.
Preferred Qualifications
  • Experience with modern ASR architectures such as Whisper, Conformer, wav2vec 2.0, DeepSpeech, or RNN-Transducer (RNN-T).
  • Experience deploying speech recognition models using ONNX Runtime, TensorRT, NVIDIA Triton Inference Server, or TorchServe.
  • Knowledge of multilingual and low-resource language speech recognition.
  • Experience with streaming speech recognition and real-time inference.
  • Familiarity with speech enhancement, voice activity detection (VAD), speaker diarization, and keyword spotting.
  • Experience with MLOps tools such as MLflow, Kubeflow, or cloud AI platforms.
  • Knowledge of Large Language Models (LLMs) for speech understanding and conversational AI.
Technical Skills
  • Python
  • PyTorch
  • TensorFlow
  • Hugging Face Transformers
  • SpeechBrain
  • ESPnet
  • Kaldi
  • torchaudio
  • librosa
  • Whisper
  • wav2vec 2.0
  • Conformer
  • RNN-T
  • ONNX Runtime
  • TensorRT
  • NVIDIA Triton Inference Server
  • TorchServe
  • Docker
  • Git
  • Linux
  • AWS / Azure / Google Cloud Platform
Soft Skills
  • Strong analytical and problem-solving skills
  • Excellent communication and collaboration
  • Attention to detail
  • Ability to work with cross-functional teams
  • Continuous learning mindset
  • Strong documentation and experimentation practices
Nice to Have
  • Experience with speech synthesis (Text-to-Speech) or conversational AI platforms
  • Knowledge of multilingual ASR evaluation and benchmarking
  • Experience with edge AI deployment for speech applications
  • Familiarity with model compression, quantization, and inference optimization
  • Publications or contributions in speech AI, ASR, or related open-source projects
Key Performance Indicators (KPIs)
  • Word Error Rate (WER) and Character Error Rate (CER)
  • Model inference latency and throughput
  • Speech recognition accuracy across languages and accents
  • Production model availability and reliability
  • Improvement in recognition quality over baseline models
  • Successful deployment and adoption of ASR features
  • Reduction in production defects and model regressions

Location
Hybrid / Remote / On-site (as applicable)
Employment Type
Full-time