1

Machine Learning Speech Jobs (NOW HIRING)

A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...

The Machine Learning Engineer will design and develop scalable training pipelines for multimodal AI ... understanding, speech-audio modeling) and dataset optimization for model training. • Solid ...

A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...

A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...

About the Role We're hiring our first Machine Learning Engineer in the United States, a ... Familiarity with multimodal data processing (e.g., text-image pairing, video understanding, speech ...

About the Role We're hiring our first Machine Learning Engineer in the United States, a ... Familiarity with multimodal data processing (e.g., text-image pairing, video understanding, speech ...

A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...

Machine Learning Engineer

Mountain View, CA · On-site +1

$196K - $221K/yr

As a Machine Learning Engineer, you'll bring your strong software engineering mindset to machine ... Is comfortable working with large-scale speech and conversational datasets, including data ...

Sr. Machine Learning Engineer, Siri Speech

Cupertino, CA · On-site

$151K - $199K/yr

... machine learning frameworks such as JAX and/or PyTorch Proficient programming skills in Python Preferred Qualifications Experience in reinforcement learning Experience with Speech LLMs or other ...

A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...

A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...

Bachelor's, Master's, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent research experience Preferred QualificationsFor Speech Researchers * Deep experience ...

Showing results 21-40

Machine Learning Speech information

What is a machine learning speech engineer?

A Machine Learning Speech Engineer is a professional who develops algorithms and models that enable computers to understand, process, and generate human speech. They work on tasks such as speech recognition, speech synthesis, speaker identification, and natural language understanding, often utilizing deep learning and other machine learning techniques. Their work is crucial for applications like virtual assistants, transcription services, and voice-controlled devices. These engineers typically have a background in computer science, linguistics, signal processing, and machine learning.

What are the key skills and qualifications needed to thrive as a machine learning speech engineer, and why are they important?

To thrive as a Machine Learning Speech Engineer, you need a solid background in computer science, signal processing, and machine learning, often supported by an advanced degree in a related field. Familiarity with tools like TensorFlow, PyTorch, Kaldi, and experience with speech recognition or natural language processing systems are typically required. Strong problem-solving skills, collaboration, and effective communication help in translating complex research into practical speech solutions. These competencies are vital for developing accurate and efficient speech technologies that meet user and business needs.

What are the typical collaboration opportunities for a machine learning speech engineer within a company?

Machine Learning Speech engineers frequently collaborate with cross-functional teams, including data scientists, software developers, product managers, and linguists. They work together to design, train, and deploy speech recognition or synthesis models, ensuring alignment with the company’s product goals and user needs. Collaboration is essential for integrating speech technologies into larger systems, troubleshooting issues, and refining models based on real-world feedback. Regular communication and teamwork help drive innovation and ensure the speech solutions are robust and user-friendly.

What is the difference between Machine Learning Speech vs Speech Recognition Engineer?

AspectMachine Learning SpeechSpeech Recognition Engineer
Required CredentialsDegree in Computer Science, Data Science, or related fields; knowledge of ML frameworksDegree in Electrical Engineering, Computer Science; experience with speech processing tools
Work EnvironmentResearch labs, tech companies, AI startupsTech companies, voice tech firms, R&D departments
Industry UsageDevelops models for speech understanding, synthesis, and processingBuilds and optimizes speech recognition systems and algorithms

Machine Learning Speech focuses on developing models for understanding and generating speech, often involving deep learning techniques. Speech Recognition Engineers specialize in creating systems that convert spoken language into text. While both roles require knowledge of speech technologies, Machine Learning Speech emphasizes model development, whereas Speech Recognition Engineers focus on system implementation and optimization.

Infographic showing various Machine Learning Speech job openings in the United States as of September 2026, with employment types broken down into 10% Internship, 70% Full Time, 10% Part Time, and 10% Contract. Highlights an 90% In-person, and 10% Remote job distribution.

Artificial Intelligence Researcher

San Jose, CA • On-site

Kotoba
Translation Services • 1 - 10 employees

Other

Posted 21 days ago


Job description

Kotoba's speech models are licensed to Fortune 50 companies and US big tech, and power an app reaching 2,000–3,000 new users a day. We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI.


Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems — models that don't just understand and generate high-quality speech, but hold the flow of a conversation: turn-taking, interruptions, overlapping speech, backchannels, response timing, prosody, and latency. Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment.


■ About Kotoba

Kotoba is a generative AI company on a mission to become the default for voice AI in East Asia. At our core is a low-latency, high-accuracy speech translation model that connects conversations so naturally it feels as though both speakers share the same language, supporting Japanese, English, Korean, Chinese, Spanish, and other major language pairs. We also build ultra-low-latency speech-to-text and text-to-speech models that run everywhere from the data center to edge devices, and we license this foundational technology to Fortune 50 companies and major US tech firms. We work from two hubs: Tokyo and San Francisco.

Our own product, the Kotoba app, is available on iOS and Android. Since launch it has grown to a steady 2,000–3,000 new downloads per day and reached No. 1 in its App Store and Google Play category, ahead of the likes of Google Translate. Enterprise adoption is accelerating in Japan, and the app has supported nearly 100 live events including SusHi Tech Tokyo.

Kotoba was founded in 2023 by two Japanese generative AI researchers with PhDs from top US universities. We've raised over ¥3 billion (roughly US$23M) from prominent VCs in Japan and the US — including Kindred Ventures and Globis Capital Partners — and from the corporate venture arms of leading US and Japanese enterprises. We also receive strong government support in Japan for AI model training.


■ What you'll do

  • Define and execute research projects for next-generation voice AI across speech-to-speech, speech-to-text, and text-to-speech systems
  • Develop full-duplex conversational models that listen and speak simultaneously while handling turn-taking, interruptions, overlapping speech, backchannels, and end-of-turn prediction
  • Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models
  • Conduct multilingual and cross-lingual research, particularly for Japanese, Korean, Chinese, English, and other languages central to our products
  • Explore architectures that orchestrate speech, language, reasoning, retrieval, and tool-use models behind a unified real-time voice interface
  • Build and scale model training and inference pipelines on distributed GPU infrastructure, optimizing models for low-latency deployment
  • Work with research, product, and infrastructure engineers to move promising research into our applications, APIs, SDKs, and customer projects


■ What we're looking for

Required

  • A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field
  • A strong research track record, demonstrated through publications at leading conferences or journals in machine learning, speech, NLP, or related areas
  • Deep expertise in at least one relevant area: speech-to-speech modeling, speech translation, spoken dialogue systems, speech recognition, speech generation, multimodal foundation models, large language models, or AI model orchestration
  • Hands-on experience designing, implementing, training, and evaluating modern neural models in PyTorch or JAX
  • Strong knowledge of modern speech and language architectures, including transformers, streaming models, autoregressive and non-autoregressive models, and foundation-model training
  • The ability to formulate original research questions, design rigorous experiments, analyze results critically, and turn promising ideas into working systems
  • Familiarity with large-scale model training, inference, data pipelines, distributed computing, and GPU-based experimentation
  • Strong written and verbal communication, including professional proficiency in English


Preferred

  • Research experience in full-duplex speech-to-speech, speech recognition, or speech generation — particularly turn-taking, interruptions, backchannels, dialogue timing, or conversational fluency
  • Research experience involving Japanese, Korean, Chinese, or other East Asian languages, including multilingual or cross-lingual modeling
  • Knowledge of audio tokenization, neural audio codecs, streaming speech recognition, streaming speech generation, or low-latency speech architectures
  • Experience with distributed training and efficient inference for large speech, language, or multimodal models
  • Research experience with systems that orchestrate multiple models, agents, retrieval components, reasoning modules, or external tools
  • Previous experience at an industrial research lab, major AI organization, technology company, or research-driven startup, particularly transferring research into production
  • A record of open-source contributions


■ Location

San Francisco