1

Speech Technology Jobs (NOW HIRING)

Sr. AI Engineer (Speech)

San Francisco, CA · On-site +1

$144K - $190K/yr

You'll bridge ML science and production engineering: whether you are strongest in research, systems, or both, you'll help turn advances in speech technology into reliable improvements for customers.

New

Showing results 41-60

Speech Technology information

See salary details

$15

$43

$69

How much do speech technology jobs pay per hour?

As of Sep 6, 2026, the average hourly pay for speech technology in the United States is $43.92, according to ZipRecruiter salary data. Most workers in this role earn between $36.06 and $51.68 per hour, depending on experience, location, and employer.

What is a speech technology?

A Speech Technology job involves developing and improving systems that enable machines to process, recognize, and synthesize human speech. Professionals in this field work with technologies like automatic speech recognition (ASR), text-to-speech (TTS), and natural language processing (NLP) to enhance voice-based applications. They collaborate with engineers, linguists, and data scientists to create efficient and accurate speech-enabled systems for industries such as AI assistants, customer service automation, and accessibility solutions.

What does a speech technology professional do?

In speech technology roles, professionals commonly work on projects such as developing and improving speech recognition systems, designing voice user interfaces, and implementing natural language processing solutions for products like virtual assistants or automated transcription tools. Day-to-day responsibilities often involve collaborating with software engineers, data scientists, and UX designers, as well as testing and optimizing algorithms using large voice datasets. You may also participate in troubleshooting, model evaluation, and refining features to meet client or product needs. This collaborative, innovative work environment offers an excellent opportunity to contribute to cutting-edge advancements in human-computer interaction.

What are the key skills and qualifications needed to thrive in speech technology?

To excel in Speech Technology, a strong background in computer science, linguistics, and machine learning—often supported by a relevant degree—is essential. Familiarity with speech recognition frameworks, natural language processing (NLP) libraries, and programming languages like Python or C++ is typically required. Strong analytical thinking, teamwork, and clear communication skills help differentiate top professionals in this field. These abilities enable the development and refinement of advanced speech systems that work accurately and efficiently for end-users.

More about Speech Technology jobs

What cities are hiring for Speech Technology jobs?

Cities with the most Speech Technology job openings:

What are the most commonly searched types of Speech Technology jobs?

The most popular types of Speech Technology jobs are:

What states have the most Speech Technology jobs?

States with the most job openings for Speech Technology jobs include:

Infographic showing various Speech Technology job openings in the United States as of August 2026, with employment types broken down into 1% As Needed, 80% Full Time, 16% Part Time, and 3% Contract. Highlights an 84% Physical, 3% Hybrid, and 13% Remote job distribution, with an average salary of $91,346 per year, or $43.9 per hour.

Sr. AI Engineer (Speech)

Dialpad

San Francisco, CA • On-site, Remote

$144K - $190K/yr

Full-time

Posted yesterday

New


Job description

About Dialpad
Dialpad is the AI platform for customer experience, built to resolve customer problems in real time across voice and digital. Our AI agents learn from your best human agents and improve with every interaction, helping organizations understand their customers, deliver better experiences, increase operational efficiencies, and build a lasting competitive advantage.

Unlike legacy systems built to route and answer, or standalone agentic bot vendors built to deflect, Dialpad was built to resolve. Our AI agents and human agents operate on a single platform with shared context, allowing Agentic AI to resolve issues, advance deals, and eliminate busywork through automation while seamlessly handing conversations to humans when needed, with full context preserved.

Market-leading brands, including Randstad, Motorola Solutions, Netflix, the San Diego Padres, the Colorado Rockies Baseball Club, and Cal Athletics, trust Dialpad. Dialpad is backed by Andreessen Horowitz, GV, ICONIQ Capital, and T-Mobile.

Being a Dialer
At Dialpad, AI isn't just a feature; it's how our teams do their best work every day. We put powerful AI tools in every employee's hands so they can move faster, think bigger, and achieve more.

We believe every conversation matters. And we've built the platform that turns those conversations into insight and action, for our customers and ourselves.

We look for people who are intensely curious and hold themselves to a high bar. Our ambition is significant, and achieving it requires a team that operates at the highest level. We seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic.

Your Role

As a Sr. AI Engineer: Speech, you'll be a senior technical leader on our Speech Team, shaping the models and systems that make Dialpad's next-generation AI voice agents accurate, natural, responsive, and trustworthy. You'll guide the team on state-of-the-art speech models, decoders, and evaluation approaches, while making hands-on contributions across speech recognition, enhancement, audio intelligence, and real-time inference. You'll bridge ML science and production engineering: whether you are strongest in research, systems, or both, you'll help turn advances in speech technology into reliable improvements for customers. This role offers broad ownership, meaningful technical influence, and the opportunity to mentor engineers while driving the next generation of Dialpad's voice experience.

This position reports to our Sr. Manager, AI Speech, and offers the opportunity to be based in our US or Canada Hub locations or work remotely.

What You'll Do

  • Technical Leadership & Direction: Set technical direction for speech-model and decoder strategy across the team, guiding evaluation of state-of-the-art approaches and turning the strongest ideas into reliable improvements for our voice agents.
  • Speech Model Development: Lead research, adaptation, and implementation of ASR/STT, speech-enhancement, and related audio models that improve recognition robustness, clarity, and performance in real-world calls.
  • Decoders & Real-Time Pipeline: Improve the speech pipeline end to end, including decoding, endpointing, turn detection, streaming behavior, and the handoff between ASR, LLM, and TTS, with a focus on natural interactions and low latency.
  • Research & Evaluation: Benchmark, prototype, fine-tune, distill, or otherwise adapt models and algorithms when doing so can create a meaningful advantage in voice-agent quality, latency, cost, or reliability.
  • Production ML & Backend: Design and ship production-grade services and inference components for real-time speech, partnering with platform and backend engineers to make model improvements observable, scalable, and maintainable.
  • Cross-Functional Leadership & Mentorship: Partner with speech, NLP, telephony, platform, product, and infrastructure engineers to shape the roadmap, lead technical reviews, mentor teammates, and translate speech advances into measurable gains for voice agents.

Skills You'll Bring

  • Speech ML & Software Engineering: Strong Python programming skills and experience with deep learning frameworks such as PyTorch, plus the ability to work effectively in production codebases. Candidates may come from a research-heavy ML background, a backend/inference engineering background, or a combination of both.
  • Speech & Audio Expertise: 5+ years of experience in speech ML, speech recognition, speech enhancement, audio AI, or a closely related field, with hands-on experience improving models or systems used in real-world applications.
  • SOTA Models & Decoders: Deep understanding of modern ASR/STT architectures and decoding techniques, with the ability to evaluate, adapt, and explain trade-offs among accuracy, robustness, streaming behavior, latency, and compute cost.
  • Research & Problem Solving: A track record of taking ideas from papers, experiments, or emerging models and turning them into measurable product improvements through disciplined prototyping, analysis, and iteration.
  • Production Systems: Experience building, deploying, and operating low-latency, production-grade ML or backend services in a cloud environment; experience with streaming inference, observability, and GCP is a plus.
  • Senior Technical Leadership: Demonstrated ability to set technical direction, communicate clearly across disciplines, mentor engineers, and make pragmatic decisions that balance model quality, reliability, latency, and customer impact.

For exceptional talent based in California, the target base salary range for this position is posted below. Our salary ranges are determined by role, level, and location. The range displayed on each job posting reflects the target range for new hire salaries for the position. Within the range, individual pay is determined by work location and additional factors, including job-related skills, experience, and relevant education or training. Your recruiter can share more about the specific salary range for your preferred location during the hiring process. Please note that the compensation details listed in US role postings reflect the base salary only, and do not include bonus, equity, or benefits.

California Salary Range
$224,500—$256,000 USD

Why Join Dialpad

  • Work at the center of the AI transformation in business communications
  • Build and ship agentic AI products that are redefining how companies operate
  • Join a team where AI amplifies every employee's impact
  • Competitive salary, comprehensive benefits, and real opportunities for growth

We believe in investing in our people. Dialpad offers competitive benefits and perks, cutting-edge AI tools, and a robust training program that help you reach your full potential. We have designed our offices to be inclusive, offering a vibrant environment to cultivate collaboration and connection. Our exceptional culture, repeatedly recognized as a Great Place to Work, ensures that every employee feels valued and empowered to contribute to our collective success.

Don't meet every single requirement? If you're excited about this role and possess the fundamental traits, drive, and strong ambition we seek, but your experience doesn't meet every qualification, we encourage you to apply. 

 Dialpad is an equal-opportunity employer. We are dedicated to creating a community of inclusion and an environment free from discrimination or harassment.