1

Audio Machine Learning Intern Jobs in Mountain View, CA

Machine Learning Research Intern, Audio As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or ...

New

To that end, there are three major components with which an intern should expect to engage. Modeling Understanding how to frame business problems as data science problems Navigating the full data ...

To that end, there are three major components with which an intern should expect to engage. Modeling Understanding how to frame business problems as data science problems Navigating the full data ...

next page

Showing results 1-20

Audio Machine Learning Intern information

See Mountain View, CA salary details

$30.1K

$50.2K

$103.8K

How much do audio machine learning intern jobs pay per year?

As of Aug 30, 2026, the average yearly pay for audio machine learning intern in Mountain View, CA is $50,235.00, according to ZipRecruiter salary data. Most workers in this role earn between $38,300.00 and $54,300.00 per year, depending on experience, location, and employer.

What does an audio machine learning intern do?

An Audio Machine Learning Intern assists in developing and improving machine learning models that process and analyze audio data. Their tasks may include data preprocessing, feature extraction, model training, and evaluation for applications like speech recognition, sound classification, or music analysis. Interns often collaborate with engineers and researchers to experiment with new algorithms and optimize audio-based AI systems. This role provides hands-on experience in both audio signal processing and machine learning techniques.

What types of projects can an audio machine learning intern expect to work on during their internship?

As an Audio Machine Learning Intern, you can expect to be involved in projects such as developing and fine-tuning audio classification models, working on speech recognition algorithms, or improving the accuracy of sound event detection systems. You may also assist with the collection and preprocessing of audio datasets, as well as support model evaluation and optimization. Collaboration with data scientists, audio engineers, and software developers is common, offering a hands-on learning environment and exposure to end-to-end machine learning workflows in the audio domain.

What are the key skills and qualifications needed to thrive as an audio machine learning intern, and why are they important?

To thrive as an Audio Machine Learning Intern, you need a solid background in signal processing, machine learning fundamentals, and programming skills, often supported by coursework or research in computer science or electrical engineering. Familiarity with Python, TensorFlow or PyTorch, and audio processing libraries like Librosa is typically required. Creativity, problem-solving abilities, and strong collaboration skills help you stand out in this role. These skills are crucial for developing innovative audio solutions, interpreting complex data, and working effectively within research or product teams.

What is the difference between Audio Machine Learning Intern vs Audio Data Analyst?

AspectAudio Machine Learning InternAudio Data Analyst
Required CredentialsTypically pursuing or recent graduate in Computer Science, Data Science, or related fieldsDegree in Data Analysis, Statistics, or related fields; may have certifications in data tools
Work EnvironmentResearch labs, tech companies, or startups focusing on AI and audio techData-driven departments within media, entertainment, or tech companies
Employer & Industry UsageUsed in AI development, research projects, and product innovationUsed for analyzing audio data, improving user experience, and reporting

The Audio Machine Learning Intern focuses on developing models and algorithms for audio data, often in research or development settings. In contrast, the Audio Data Analyst primarily interprets audio data to generate insights and support decision-making. Both roles require familiarity with audio data, but the intern role emphasizes machine learning skills, while the analyst role centers on data analysis and reporting.

What are popular job titles related to Audio Machine Learning Intern jobs in Mountain View, CA?

For Audio Machine Learning Intern jobs in Mountain View, CA, the most frequently searched job titles are:

What job categories do people searching Audio Machine Learning Intern jobs in Mountain View, CA look for?

The top searched job categories for Audio Machine Learning Intern jobs in Mountain View, CA are:

What cities near Mountain View, CA are hiring for Audio Machine Learning Intern jobs?

Cities near Mountain View, CA with the most Audio Machine Learning Intern job openings:

Machine Learning Intern

San Francisco, CA โ€ข On-site

Bland
Clean Energy Semiconductors Manufacturingย โ€ขย 51 - 200 employees

Full-time

Posted 2 days ago

New


Job description

The Role: Machine Learning Research Intern, Audio
As a Research Intern at Bland, you will own a focused research project across our voice stack: speech-to-text, large language models, neural audio codecs, or text-to-speech. You will work alongside our research team on the same problems they are working on, not on a side track built to keep interns busy.
We scope internships around a single meaningful question that can be answered in the time you have. The goal is a result worth shipping, publishing, or both. Interns here regularly see their work reach production systems handling millions of calls.
What You Will Do
Own a research question end to end
  • Take one well-scoped problem from literature review through implementation, experimentation, and results.
  • Design ablations that isolate what actually caused an improvement.
  • Present your findings to the research team and defend the methodology.

Work on real systems
  • Train and evaluate models on large-scale, real-world telephony audio, including the accents, noise, and artifacts that make production speech hard.
  • Use our distributed GPU infrastructure rather than toy-scale setups.
  • Where the result warrants it, work with engineers to move it toward production.

Choose your depth
Depending on your background and interests, your project may focus on:
  • Expressive and controllable text-to-speech, including prosody and emotion modeling
  • Neural audio codecs and discrete or continuous speech representations
  • ASR robustness for telephony, accents, and code switching
  • Real-time and streaming inference under latency constraints
  • Full-duplex conversation and turn-taking dynamics
What Makes You a Great Fit
Research foundations
  • Currently pursuing a MS or PhD in ML, CS, EE, or a related field, or equivalent research experience.
  • Comfortable reading a paper and reimplementing it without hand-holding.
  • Experience with self-supervised, generative, or multimodal modeling.

Audio or speech grounding
  • Hands-on work with speech or audio models, whether TTS, ASR, codecs, or audio representation learning.
  • Strong intuition for audio quality and what makes synthetic speech sound wrong.
  • Prior publications or open source contributions in speech or language AI are a strong signal, though not required.

Engineering ability
  • Fluent in PyTorch and comfortable in a real codebase.
  • Able to run your own experiments on GPU clusters without waiting to be unblocked.
How You Show Up
  • You identify the single experiment that validates an idea in days, not months.
  • You measure everything and let data drive decisions.
  • You are honest about negative results, because they are how we narrow the search.
  • You are obsessed with making voice agents sound truly human.
  • You use AI tools aggressively to amplify your own impact.
Benefits
  • Competitive intern compensation
  • Mentorship from researchers working on frontier voice AI
  • Every tool you need to succeed
  • Beautiful office in Levi's Plaza, SF with rooftop views
  • A real shot at a return offer