1

Research Scientist Audio Jobs (NOW HIRING)

Research & Experimentation - Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems. * Model Training - Rapidly ...

We are seeking Research Scientists (across all levels of seniority) to join our Artist-First AI ... What You'll Do Conduct groundbreaking research in generative audio using diffusion or flow matching ...

We are seeking Research Scientists (across all levels of seniority) to join our Artist-First AI ... What You'll Do Conduct groundbreaking research in generative audio using diffusion or flow matching ...

They are seeking a Research Scientist to develop and fine-tune models for video and audio data processing and enhancement, as well as conduct data-oriented research to improve multimodal quality.

About The Role As a Research Scientist at Phonic, you'll drive original research that pushes the ... Deep expertise in at least one of: speech synthesis, ASR, audio modeling, language modeling, or ...

next page

Showing results 1-20

Research Scientist Audio information

See salary details

$50.5K

$130.1K

$174K

How much do research scientist audio jobs pay per year?

As of Aug 11, 2026, the average yearly pay for research scientist audio in the United States is $130,117.00, according to ZipRecruiter salary data. Most workers in this role earn between $107,500.00 and $173,000.00 per year, depending on experience, location, and employer.

What does a research scientist in audio do?

A Research Scientist in Audio focuses on advancing the science and technology related to sound and audio processing. Their work often involves developing algorithms for speech recognition, sound classification, audio enhancement, and spatial audio technologies. They may conduct experiments, analyze data, and collaborate with engineers to create innovative solutions for applications such as virtual assistants, music technology, or hearing aids. This role typically requires a strong background in signal processing, machine learning, and acoustics. Research Scientists in Audio also publish their findings and contribute to the broader scientific community.

What are the key skills and qualifications needed to thrive as a research scientist in audio?

To thrive as a Research Scientist in Audio, you need advanced knowledge in signal processing, machine learning, and acoustics, usually supported by a PhD or equivalent experience in computer science, electrical engineering, or a related field. Familiarity with programming languages like Python or MATLAB, experience with deep learning frameworks (e.g., TensorFlow, PyTorch), and knowledge of audio analysis tools are essential. Strong analytical thinking, creativity, and effective communication skills distinguish top performers in this field. These competencies enable the development of innovative audio technologies and effective collaboration within multidisciplinary research teams.

How does a research scientist in audio typically collaborate with cross-functional teams during product development?

As a Research Scientist in audio, you will frequently work alongside software engineers, product managers, and UX designers to integrate cutting-edge audio algorithms into products. Collaboration often involves presenting research findings, prototyping solutions, and iterating based on feedback from both technical and non-technical team members. This role requires strong communication skills to bridge the gap between complex research concepts and practical implementation, ensuring that new audio technologies align with user needs and project goals.

What is the difference between Research Scientist Audio vs Audio Engineer?

AspectResearch Scientist AudioAudio Engineer
Required CredentialsMaster's or PhD in Acoustics, Audio Engineering, or related fieldsBachelor's degree in Audio Engineering, Sound Design, or related fields
Work EnvironmentResearch labs, universities, R&D departmentsRecording studios, live events, broadcasting facilities
Employer & Industry UsageAcademic institutions, tech companies, research organizationsMusic, film, broadcasting, live sound production

Research Scientist Audio focuses on developing new audio technologies and conducting experiments, often within academic or research settings. In contrast, Audio Engineers primarily operate sound equipment for recording, mixing, and live sound. While both roles require a strong understanding of audio principles, Research Scientists typically hold advanced degrees and work in R&D environments, whereas Audio Engineers often have practical experience and work directly with sound production equipment.

More about Research Scientist Audio jobs
What cities are hiring for Research Scientist Audio jobs? Cities with the most Research Scientist Audio job openings:
What states have the most Research Scientist Audio jobs? States with the most job openings for Research Scientist Audio jobs include:
Infographic showing various Research Scientist Audio job openings in the United States as of August 2026, with employment types broken down into 63% Full Time, 33% Part Time, and 4% Contract. Highlights an 100% In-person job distribution, with an average salary of $130,117 per year, or $62.6 per hour.

Research Scientist - Audio

Retell AI

San Francisco, CA • On-site

$225K - $400K/yr

Full-time

Medical, Dental, Vision

Posted 7 days ago


Job description

ABOUT RETELL AI
Retell AI is using first-principles thinking to reimagine the call center with cutting-edge voice AI. Thousands of companies now use Retell's AI voice agents to handle sales, support, and logistics calls that once required large teams of human agents. Backed by Y Combinator, Alt Capital, and other leading investors, we've scaled to $80M in ARR with a team of 50, up from $5M at the start of 2025, and are now valued at over $1.5B.
Our vision for 2026 is to build a modern CX platform where entire contact centers are powered by AI. Instead of basic automation that needs constant human tuning, we're creating intelligent AI "workers" that act as frontline agents, QA analysts, and managers, continuously executing, monitoring, and improving every customer interaction.
We're growing fast and looking for ambitious builders who want to tackle hard technical problems, move quickly, and have a real impact on one of the fastest-growing voice AI companies in the world.
Let's build the future together.
Recent recognition:
  • No. 1 Best Places to Work in the Bay Area, San Francisco Business Times 2026
  • Top 50 AI Apps, a16z (2025)
  • #3 Fastest-Growing Software Company, G2 Best Software Awards 2026
  • Best Agentic AI Software, G2 Best Software Awards 2026
  • #4 Fastest-Growing Software Vendor, Brex Benchmark 2025
  • Enterprise Tech 30 Class of 2026, Nasdaq & Wing VC
  • Top-Ranked Startup, Lean AI Leaderboard
  • Backed by Y Combinator

ABOUT THE ROLE
Retell AI transforms customer experience with voice AI for enterprises, including customers like CVS/Aetna, American Airlines, Lenovo, and Grab. We have more customer stories than we can tell!
This is a research-driven, high-impact role for ML researchers who want to push the boundaries of real-time AI. As a Founding Machine Learning Research Engineer at Retell, you'll focus on advancing model capabilities for human-like voice agents operating in complex, real-world environments.
You'll explore new approaches across LLMs and audio models, design novel evaluation methods, and prototype systems that improve reasoning, latency, and conversational quality. Your work will directly influence production systems, bridging cutting-edge research with real-world deployment.
If you're excited about solving open-ended ML problems, experimenting rapidly, and shaping how voice AI systems think and perform, this is a unique opportunity to do so at scale.
KEY RESPONSIBILITIES
  • Research & Experimentation - Explore and develop new techniques across LLMs and audio models to improve reasoning, latency, and conversational quality in real-time systems.
  • Model Training - Rapidly build and iterate on models and pipelines, turning research ideas into working prototypes. Innovate on paradigms, training methods, and inference.
  • Evaluation & Benchmarking - Design novel evaluation frameworks, datasets, and metrics to measure performance on complex, real-world voice tasks.
  • Bridge Research to Production - Collaborate closely with engineering to translate research insights into deployable systems.
  • Human Feedback Loops - Develop methods to incorporate human evaluation into model improvement, especially for subjective conversational quality.
  • Advance the Frontier - Stay at the cutting edge of ML research and bring new ideas into Retell's product and infrastructure.

REQUIRED
  • Strong ML Research Background - You've worked on advanced ML problems (like LLM pre-training and post-training, transcription model training, TTS, or multimodal systems), either in industry or academia.
  • Deep Technical Foundation - Comfortable with PyTorch, model architectures, and the math behind modern machine learning.
  • Top Academic Background - Master's degree in CS, ML, AI or related field required; PhD preferred. Equivalent research-level engineering experience also considered.

YOU MIGHT THRIVE IF YOU
  • Published or Awarded - First/co-author publications at top-tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, etc.) or notable competition awards are a strong plus.
  • Experimental Mindset - You enjoy exploring open-ended problems and iterating quickly on ideas.
  • Bridge Theory & Practice - You can translate research into systems that work in real-world environments.
  • Startup-Ready - You thrive in fast-paced environments with high ownership and ambiguity.
  • Collaborative & Clear Communicator - You can explain complex ideas and work cross-functionally to drive impact.

JOB DETAILS
  • Cash: $225,000 - $400,000 base salary
  • Equity: Offers Equity
  • Location: Redwood City, CA, US (100% Relocation Provided)
  • US Visas: Retell AI is open to sponsoring work authorization for qualified candidates, including H1B/H-1B, TN, L-1, E-3, F-1 (OPT/CPT).

OTHER BENEFITS
  • 100% coverage for medical, dental, and vision insurance
  • $70/day DoorDash credit for unlimited meals and snacks
  • $200/month wellness reimbursement
  • $300/month commuter reimbursement
  • $75/month phone bill reimbursement
  • $50/month internet reimbursement
COMPENSATION PHILOSOPHY
  • Best Offer Upfront: Choose from three cash-equity balance options, no negotiation needed
  • Top 1% Talent: Above-market pay (top 5 percentile)
  • High Ownership: Small teams, >$1M revenue/employee, significant equity
  • Performance-Based: Offers tied to interview performance, not past salaries
INTERVIEW PROCESS
  • Talent Screen (15min): chat with our recruiter to get a better sense of the role, the team, and what it's like to work here.
  • Technical Interview (45 min): LLM theory specific coding Interview (PyTorch)
  • Technical Interview (45 min): Live Practical Systems Design and Coding Interview.
  • Onsite/Virtual Interviews (3 hrs): Hosted in our office if located in the Bay Area or virtual, with three rounds:
    • ML System Design
    • ML Question Deep Dive
    • Backend + AI Practical