A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
The Machine Learning Engineer will design and develop scalable training pipelines for multimodal AI ... understanding, speech-audio modeling) and dataset optimization for model training. • Solid ...
The Machine Learning Engineer will design and develop scalable training pipelines for multimodal AI ... understanding, speech-audio modeling) and dataset optimization for model training. • Solid ...
Required : • Deep expertise in areas such as machine learning, speech recognition, natural language processing, computer vision, knowledge acquisition, etc. • Ability to assess the feasibility ...
Required : • Deep expertise in areas such as machine learning, speech recognition, natural language processing, computer vision, knowledge acquisition, etc. • Ability to assess the feasibility ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
Machine Learning Engineer
Mountain View, CA · On-site
$175K - $275K/yr
About the Role We're hiring our first Machine Learning Engineer in the United States, a ... Familiarity with multimodal data processing (e.g., text-image pairing, video understanding, speech ...
Machine Learning Engineer
Mountain View, CA · On-site
$175K - $275K/yr
About the Role We're hiring our first Machine Learning Engineer in the United States, a ... Familiarity with multimodal data processing (e.g., text-image pairing, video understanding, speech ...
Machine Learning Engineer
Mountain View, CA · On-site
$175K - $275K/yr
About the Role We're hiring our first Machine Learning Engineer in the United States, a ... Familiarity with multimodal data processing (e.g., text-image pairing, video understanding, speech ...
Machine Learning Engineer
Mountain View, CA · On-site
$175K - $275K/yr
About the Role We're hiring our first Machine Learning Engineer in the United States, a ... Familiarity with multimodal data processing (e.g., text-image pairing, video understanding, speech ...
Sr. Machine Learning Engineer, Siri Speech
Cupertino, CA · On-site
$184K - $324K/yr
Sr. Machine Learning Engineer, Siri Speech Cupertino, California, United States Machine Learning and AI Join the team redefining what a deeply personal and integrated assistant can be.As part of the ...
Sr. Machine Learning Engineer, Siri Speech
Cupertino, CA · On-site
$184K - $324K/yr
Sr. Machine Learning Engineer, Siri Speech Cupertino, California, United States Machine Learning and AI Join the team redefining what a deeply personal and integrated assistant can be.As part of the ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
The Machine Learning Intern will work alongside senior ML engineers to help build, evaluate, and ... learning for audio or speech (coursework, research, or project experience with CNNs/RNNs on ...
New
Quick apply
The Machine Learning Intern will work alongside senior ML engineers to help build, evaluate, and ... learning for audio or speech (coursework, research, or project experience with CNNs/RNNs on ...
New
The Machine Learning Intern will work alongside senior ML engineers to help build, evaluate, and ... learning for audio or speech (coursework, research, or project experience with CNNs/RNNs on ...
New
The Machine Learning Intern will work alongside senior ML engineers to help build, evaluate, and ... learning for audio or speech (coursework, research, or project experience with CNNs/RNNs on ...
New
Machine Learning Engineer
Mountain View, CA · On-site +1
$196K - $221K/yr
As a Machine Learning Engineer, you'll bring your strong software engineering mindset to machine ... Is comfortable working with large-scale speech and conversational datasets, including data ...
Machine Learning Engineer
Mountain View, CA · On-site +1
$196K - $221K/yr
As a Machine Learning Engineer, you'll bring your strong software engineering mindset to machine ... Is comfortable working with large-scale speech and conversational datasets, including data ...
Sr. Machine Learning Scientist, Siri Speech
Cupertino, CA · On-site
$184K - $324K/yr
... learning emerge when we address real-world problems at scale. We develop speech to speech ... machine learning training/evaluation Data-centric vision about foundation model Minimum ...
Sr. Machine Learning Scientist, Siri Speech
Cupertino, CA · On-site
$184K - $324K/yr
... learning emerge when we address real-world problems at scale. We develop speech to speech ... machine learning training/evaluation Data-centric vision about foundation model Minimum ...
The Machine Learning Intern will work alongside senior ML engineers to help build, evaluate, and ... learning for audio or speech (coursework, research, or project experience with CNNs/RNNs on ...
New
The Machine Learning Intern will work alongside senior ML engineers to help build, evaluate, and ... learning for audio or speech (coursework, research, or project experience with CNNs/RNNs on ...
New
Sr. Machine Learning Engineer, Siri Speech
Cupertino, CA · On-site
$151K - $199K/yr
... machine learning frameworks such as JAX and/or PyTorch Proficient programming skills in Python Preferred Qualifications Experience in reinforcement learning Experience with Speech LLMs or other ...
Sr. Machine Learning Engineer, Siri Speech
Cupertino, CA · On-site
$151K - $199K/yr
... machine learning frameworks such as JAX and/or PyTorch Proficient programming skills in Python Preferred Qualifications Experience in reinforcement learning Experience with Speech LLMs or other ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field * A strong research ...
Bachelor's, Master's, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent research experience Preferred QualificationsFor Speech Researchers * Deep experience ...
Bachelor's, Master's, or PhD in Computer Science, Machine Learning, or a related technical field, or equivalent research experience Preferred QualificationsFor Speech Researchers * Deep experience ...
Machine Learning Speech information
What is a machine learning speech engineer?
What are the key skills and qualifications needed to thrive as a machine learning speech engineer, and why are they important?
What are the typical collaboration opportunities for a machine learning speech engineer within a company?
What is the difference between Machine Learning Speech vs Speech Recognition Engineer?
| Aspect | Machine Learning Speech | Speech Recognition Engineer |
|---|---|---|
| Required Credentials | Degree in Computer Science, Data Science, or related fields; knowledge of ML frameworks | Degree in Electrical Engineering, Computer Science; experience with speech processing tools |
| Work Environment | Research labs, tech companies, AI startups | Tech companies, voice tech firms, R&D departments |
| Industry Usage | Develops models for speech understanding, synthesis, and processing | Builds and optimizes speech recognition systems and algorithms |
Machine Learning Speech focuses on developing models for understanding and generating speech, often involving deep learning techniques. Speech Recognition Engineers specialize in creating systems that convert spoken language into text. While both roles require knowledge of speech technologies, Machine Learning Speech emphasizes model development, whereas Speech Recognition Engineers focus on system implementation and optimization.
What other helpful pages are available for Machine Learning Speech?
Other pages related to Machine Learning Speech:

Artificial Intelligence Researcher
San Jose, CA • On-site
Other
Posted 21 days ago
Job description
Kotoba's speech models are licensed to Fortune 50 companies and US big tech, and power an app reaching 2,000–3,000 new users a day. We're hiring an AI Researcher to build the next generation of real-time, interactive voice AI.
Location: San Francisco. You'll work on full-duplex speech-to-speech, speech-to-text, and text-to-speech systems — models that don't just understand and generate high-quality speech, but hold the flow of a conversation: turn-taking, interruptions, overlapping speech, backchannels, response timing, prosody, and latency. Your work runs the full research lifecycle, from framing the question through distributed training, evaluation, and production deployment.
■ About Kotoba
Kotoba is a generative AI company on a mission to become the default for voice AI in East Asia. At our core is a low-latency, high-accuracy speech translation model that connects conversations so naturally it feels as though both speakers share the same language, supporting Japanese, English, Korean, Chinese, Spanish, and other major language pairs. We also build ultra-low-latency speech-to-text and text-to-speech models that run everywhere from the data center to edge devices, and we license this foundational technology to Fortune 50 companies and major US tech firms. We work from two hubs: Tokyo and San Francisco.
Our own product, the Kotoba app, is available on iOS and Android. Since launch it has grown to a steady 2,000–3,000 new downloads per day and reached No. 1 in its App Store and Google Play category, ahead of the likes of Google Translate. Enterprise adoption is accelerating in Japan, and the app has supported nearly 100 live events including SusHi Tech Tokyo.
Kotoba was founded in 2023 by two Japanese generative AI researchers with PhDs from top US universities. We've raised over ¥3 billion (roughly US$23M) from prominent VCs in Japan and the US — including Kindred Ventures and Globis Capital Partners — and from the corporate venture arms of leading US and Japanese enterprises. We also receive strong government support in Japan for AI model training.
■ What you'll do
- Define and execute research projects for next-generation voice AI across speech-to-speech, speech-to-text, and text-to-speech systems
- Develop full-duplex conversational models that listen and speak simultaneously while handling turn-taking, interruptions, overlapping speech, backchannels, and end-of-turn prediction
- Improve the accuracy, naturalness, expressiveness, multilingual robustness, and streaming latency of speech recognition and speech generation models
- Conduct multilingual and cross-lingual research, particularly for Japanese, Korean, Chinese, English, and other languages central to our products
- Explore architectures that orchestrate speech, language, reasoning, retrieval, and tool-use models behind a unified real-time voice interface
- Build and scale model training and inference pipelines on distributed GPU infrastructure, optimizing models for low-latency deployment
- Work with research, product, and infrastructure engineers to move promising research into our applications, APIs, SDKs, and customer projects
■ What we're looking for
Required
- A PhD or equivalent research experience in machine learning, speech processing, natural language processing, multimodal AI, human-computer interaction, or a closely related field
- A strong research track record, demonstrated through publications at leading conferences or journals in machine learning, speech, NLP, or related areas
- Deep expertise in at least one relevant area: speech-to-speech modeling, speech translation, spoken dialogue systems, speech recognition, speech generation, multimodal foundation models, large language models, or AI model orchestration
- Hands-on experience designing, implementing, training, and evaluating modern neural models in PyTorch or JAX
- Strong knowledge of modern speech and language architectures, including transformers, streaming models, autoregressive and non-autoregressive models, and foundation-model training
- The ability to formulate original research questions, design rigorous experiments, analyze results critically, and turn promising ideas into working systems
- Familiarity with large-scale model training, inference, data pipelines, distributed computing, and GPU-based experimentation
- Strong written and verbal communication, including professional proficiency in English
Preferred
- Research experience in full-duplex speech-to-speech, speech recognition, or speech generation — particularly turn-taking, interruptions, backchannels, dialogue timing, or conversational fluency
- Research experience involving Japanese, Korean, Chinese, or other East Asian languages, including multilingual or cross-lingual modeling
- Knowledge of audio tokenization, neural audio codecs, streaming speech recognition, streaming speech generation, or low-latency speech architectures
- Experience with distributed training and efficient inference for large speech, language, or multimodal models
- Research experience with systems that orchestrate multiple models, agents, retrieval components, reasoning modules, or external tools
- Previous experience at an industrial research lab, major AI organization, technology company, or research-driven startup, particularly transferring research into production
- A record of open-source contributions
■ Location
San Francisco
About Kotoba
Sourced by ZipRecruiter
Industry
Translation services
Company size
1 - 10 Employees
Headquarters location
Bethesda, MD, US