1

Speech Recognition Jobs (NOW HIRING)

Staff ML Engineer

San Francisco, CA · On-site

$180 - $240/hr

You'll work on cutting-edge problems in medical speech recognition, clinical language understanding, and agentic AI systems for healthcare. Key Responsibilities * Develop and optimize models for ...

This role involves working with state-of-the-art speech recognition, speech synthesis, voice streaming, and AI agent frameworks to create scalable voice solutions for customer service, sales, finance ...

Develop and train speech recognition models optimized for medical conversations across hundreds of specialties * Leverage Knowtex's large proprietary clinical audio dataset to train and fine-tune ...

Develop and train speech recognition models optimized for medical conversations across hundreds of specialties * Leverage Knowtex's large proprietary clinical audio dataset to train and fine-tune ...

This role involves working with state-of-the-art speech recognition, speech synthesis, voice streaming, and AI agent frameworks to create scalable voice solutions for customer service, sales, finance ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models * Conduct multilingual and cross-lingual research ...

Knowledge of speech recognition and synthesis technologies. • Computer Vision: Experience with computer vision tasks such as image classification, object detection, and segmentation. • (Retrieval ...

Showing results 21-40

Speech Recognition information

See salary details

$15

$43

$69

How much do speech recognition jobs pay per hour?

As of Sep 4, 2026, the average hourly pay for speech recognition in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is speech recognition?

A Speech Recognition job involves developing and improving systems that convert spoken language into text. Professionals in this field work with machine learning, natural language processing (NLP), and signal processing to enhance speech-to-text accuracy. They may design models, train algorithms, and fine-tune systems for applications like virtual assistants, transcription services, and accessibility tools. Strong programming skills and knowledge of AI technologies are essential for success in this role.

What does a speech recognition professional do?

On a daily basis, speech recognition professionals are typically involved in designing, training, and optimizing speech-to-text models, analyzing audio data, and troubleshooting system errors. They often collaborate with software engineers, data scientists, and product teams to integrate speech technologies into various applications. Responsibilities may also include evaluating model performance using large speech datasets, updating acoustic or language models, and keeping up to date with advancements in machine learning techniques. This blend of technical and collaborative work ensures continuous improvement and innovation in speech recognition products.

What skills and qualifications are needed for speech recognition?

To excel in Speech Recognition, you need a strong background in computational linguistics, machine learning, and signal processing, often supported by a degree in computer science, engineering, or a related field. Familiarity with tools such as Python, TensorFlow, Kaldi, and speech corpus databases, alongside relevant certifications in AI or data science, is highly valuable. Strong analytical skills, attention to detail, and the ability to collaborate effectively with cross-functional teams are important soft skills. These competencies are critical for developing, refining, and implementing accurate and efficient speech recognition systems that meet real-world needs.

More about Speech Recognition jobs

What cities are hiring for Speech Recognition jobs?

Cities with the most Speech Recognition job openings:

What are the most commonly searched types of Speech Recognition jobs?

The most popular types of Speech Recognition jobs are:

What states have the most Speech Recognition jobs?

States with the most job openings for Speech Recognition jobs include:

Infographic showing various Speech Recognition job openings in the United States as of August 2026, with employment types broken down into 91% Full Time, and 9% Part Time. Highlights an 91% In-person, and 9% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Staff ML Engineer

SupportFinity™

San Francisco, CA • On-site

$180 - $240/hr

Other

Medical, Dental, Vision, Retirement, PTO

Posted 9 days ago


Job description

About Knowtex

Knowtex is building the future of voice AI operating systems for clinicians, transforming how healthcare documentation happens at the point of care. Founded by Stanford AI scientists with deep clinical experience, we're experiencing explosive growth across both commercial health systems and federal healthcare, with our ambient documentation platform scaling rapidly to thousands of clinicians across hundreds of specialties. We're at an inflection point where cutting-edge AI meets real clinical impact, giving clinicians hours back each day to focus on what matters most - their patients.

Position Overview

We're seeking a Staff ML Engineer to advance our voice AI and clinical NLP capabilities. You'll work on cutting-edge problems in medical speech recognition, clinical language understanding, and agentic AI systems for healthcare.

Key Responsibilities
  • Develop and optimize models for medical speech recognition across 200+ specialties
  • Build clinical NLP pipelines for automated E&M coding and ICD-10 classification
  • Implement note quality evaluation systems using LLMs and clinical rubrics
  • Scale inference infrastructure using Triton Inference Server on AWS GovCloud
  • Create specialty-specific language models for gastroenterology, dermatology, and emerging markets
  • Design agentic AI systems for clinical decision support and documentation assistance
  • Optimize model performance for real-time inference with sub-200ms latency requirements
  • Collaborate with clinical teams to validate model outputs against MDM levels
  • Build evaluation frameworks for MIPS quality measures compliance
Required Qualifications
  • 5+ years experience in ML engineering with focus on NLP/speech recognition
  • Strong expertise in PyTorch or TensorFlow
  • Experience with transformer architectures and large language models
  • Proficiency in building production ML pipelines at scale
  • Understanding of model optimization techniques (quantization, distillation, pruning)
  • Experience with cloud ML platforms (AWS SageMaker, GCP Vertex AI)
  • Master's or PhD in Computer Science, ML, or related field
Preferred Qualifications
  • Healthcare or clinical NLP experience
  • Familiarity with medical terminology and clinical documentation
  • Experience with speech recognition systems (Whisper, Conformer architectures)
  • Knowledge of medical coding systems (CPT, ICD-10, SNOMED)
  • Publications in ML/NLP conferences
Benefits
  • Meaningful equity compensation
  • Unlimited PTO
  • Premium health, dental, and vision coverage
  • 401(k) plan
  • Hybrid work model: 3 days/week in our San Francisco office
About the company

Knowtex

#J-18808-Ljbffr