1

Speech Research Jobs (NOW HIRING)

Speech/Audio Centific AI ResearchAbout Centific AI Research Centific AI Research is at the forefront of developing cutting-edge AI solutions that bridge the gap between research innovation and real ...

Speech & Audio Research Engineer

San Diego, CA · On-site

$148.30 - $222.50/hr

General Summary Qualcomm's Multimedia R&D Group is seeking candidates to join its world‑class Speech team developing novel technologies enabling the future of real‑time communication and ...

We are looking for researchers with deep expertise in one of two areas: * Speech & Audio: Build state-of-the-art medical speech-to-text systems using our large proprietary dataset of real-world ...

We are looking for researchers with deep expertise in one of two areas: Speech & Audio: Build state-of-the-art medical speech-to-text systems using our large proprietary dataset of real-world ...

New

We are looking for researchers with deep expertise in one of two areas: * Speech & Audio: Build state-of-the-art medical speech-to-text systems using our large proprietary dataset of real-world ...

Speech Therapist

Fair Lawn, NJ · On-site

$65 - $72/hr

... research and best practices in speech-language pathology Job Types: Full-time Benefits: * 401(k) * 401(k) matching * Continuing education credits * Dental insurance * Health insurance * Paid time off

Speech Therapist

Fair Lawn, NJ · On-site

$65 - $73/hr

... research and best practices in speech-language pathology Job Type: Part-time Pay: $68.00 - $73.00 per hour Expected hours: 6 per week Benefits: * 401(k) * 401(k) matching * Continuing education ...

Speech Therapist

Fair Lawn, NJ · On-site

$65 - $72/hr

... research and best practices in speech-language pathology Job Types: Full-time Benefits: * 401(k) * 401(k) matching * Continuing education credits * Dental insurance * Health insurance * Paid time off

Speech Therapist

Fair Lawn, NJ · On-site

$65 - $72/hr

... research and best practices in speech-language pathology Job Types: Full-time Benefits: * 401(k) * 401(k) matching * Continuing education credits * Dental insurance * Health insurance * Paid time off

Speech Therapist

Fair Lawn, NJ · On-site

$65 - $73/hr

... research and best practices in speech-language pathology Job Type: Part-time Pay: $68.00 - $73.00 per hour Expected hours: 6 per week Benefits: * 401(k) * 401(k) matching * Continuing education ...

Speech Therapist

Fair Lawn, NJ · On-site

$65 - $73/hr

... research and best practices in speech-language pathology Job Type: Part-time Pay: $68.00 - $73.00 per hour Expected hours: 6 per week Benefits: * 401(k) * 401(k) matching * Continuing education ...

Showing results 21-40

Speech Research information

See salary details

$12

$45

$67

How much do speech research jobs pay per hour?

As of Aug 23, 2026, the average hourly pay for speech research in the United States is $45.78, according to ZipRecruiter salary data. Most workers in this role earn between $33.65 and $56.49 per hour, depending on experience, location, and employer.

What is speech research?

A Speech Research job involves studying and developing technologies related to speech processing, recognition, synthesis, and understanding. Researchers in this field work on improving voice assistants, automatic transcription, language modeling, and speech-based AI systems. The role often involves machine learning, linguistics, signal processing, and artificial intelligence to enhance human-computer interaction. Professionals in speech research may work in academia, tech companies, or research institutions to advance the capabilities of speech-driven applications.

What are the typical responsibilities of a speech researcher within a team?

Speech Researchers often work as part of interdisciplinary teams, collaborating with engineers, data scientists, and product managers to design and conduct experiments involving speech recognition, synthesis, or processing. Daily tasks may include collecting and annotating audio data, developing algorithms, analyzing linguistic features, and publishing findings in peer-reviewed forums. Progress is usually shared in regular team meetings, and researchers are encouraged to stay current with advancements in the field. This collaborative environment fosters innovation and the practical application of research in real-world speech technology products.

What are the key skills and qualifications needed to thrive in speech research, and why are they important?

To thrive in Speech Research, you typically need a solid background in linguistics, computational linguistics, or a related field, with strong analytical and experimental design skills. Familiarity with speech analysis software, programming languages such as Python, and tools like Praat or MATLAB is often essential. Exceptional problem-solving abilities, attention to detail, and collaborative communication skills help you excel in multidisciplinary research teams. These competencies enable researchers to conduct high-quality experiments and contribute innovative advances in speech technology.

More about Speech Research jobs

What cities are hiring for Speech Research jobs?

Cities with the most Speech Research job openings:

What are the most commonly searched types of Speech Research jobs?

The most popular types of Speech Research jobs are:

What states have the most Speech Research jobs?

States with the most job openings for Speech Research jobs include:

Infographic showing various Speech Research job openings in the United States as of August 2026, with employment types broken down into 1% Internship, 1% As Needed, 85% Full Time, 11% Part Time, and 2% Contract. Highlights an 84% Physical, 4% Hybrid, and 12% Remote job distribution, with an average salary of $95,216 per year, or $45.8 per hour.

AI Research Engineer- Speech 1

Centific Global Solutions, Inc.

Redmond, WA • On-site

$150 - $160/hr

Other

Posted 5 days ago


Job description

Key Responsibilities
  • Design, develop, and deploy Large Audio Language Models (LALMs) capable of native audio understanding, reasoning, and generation.
  • Build Large Audio Reasoning Models that perform complex chain-of-thought reasoning over speech and audio inputs across medical, technical, and conversational domains.
  • Contribute to Speech-to-Speech (S2S) system development, including speech understanding, dialogue management, and speech synthesis components.
  • Research and implement alignment mechanisms between speech encoders and LLM backbones using lightweight adapters, LoRA, and efficient fine-tuning strategies.
  • Design efficient speech tokenization and temporal compression techniques suitable for long‑form audio reasoning and multi‑turn spoken dialogue.
  • Build comprehensive evaluation frameworks for audio reasoning capabilities, including benchmarks for speech QA, audio understanding, and reasoning accuracy.
  • Optimize inference pipelines for low‑latency, streaming applications in speech systems.
  • Collaborate with cross‑functional teams to transfer research innovations into production systems and customer‑facing applications.
  • Contribute to technical documentation, research write‑ups, and publications at top‑tier venues (NeurIPS, ICML, ACL, Interspeech).
Minimum Qualifications
  • Master's degree (required) or Ph.D. (preferred) in Computer Science, Electrical Engineering, or a related field with a focus on speech, audio ML, or multimodal learning.
  • 2+ years of industry or applied research experience in speech/audio AI, Large Language Models, or multimodal systems.
  • Demonstrated applied research contributions through publications, patents, or shipped products in speech/audio AI or LLMs.
  • Strong proficiency in Python and PyTorch, with hands‑on experience in GPU‑accelerated training for large‑scale models.
  • Solid understanding of speech and audio signal processing, acoustic modeling, and audio representations.
  • Working knowledge of modern LLM architectures (Transformers, SSMs) and training paradigms including instruction tuning and alignment methods.
  • Familiarity with modality alignment techniques: adapter‑based integration, cross‑modal attention, or audio‑text fusion methods.
  • Strong experimentation habits: clean code, systematic ablations, reproducibility, and clear technical communication.
Preferred Qualifications
  • Publication record at top‑tier venues (NeurIPS, ICML, ICLR, ACL, Interspeech, ICASSP) in audio language models, speech reasoning, or multimodal learning.
  • Hands‑on experience building or fine‑tuning Large Audio Language Models (e.g., Qwen‑Audio, SALMONN, LTU, Gemini Audio).
  • Experience with speech representation pretraining (HuBERT, Wav2Vec2.0, Whisper, WavLM) and discrete speech tokenization.
  • Familiarity with Speech‑to‑Speech components: neural audio codecs (EnCodec, SoundStream), vocoders, or speech synthesis systems.
  • Experience with audio reasoning benchmarks (AIR‑Bench, MMAU, AudioBench) or building evaluation harnesses for audio QA.
  • Hands‑on experience with distributed training (FSDP, DeepSpeed) and inference optimization (ONNX, TensorRT, quantization).
  • Familiarity with speech frameworks such as ESPnet, SpeechBrain, NVIDIA NeMo, or Fairseq.
  • Experience with multilingual speech systems, code‑switching, or domain adaptation for specialized applications (medical, legal, technical).
  • Background in evaluating safety, bias, hallucination, or adversarial robustness in audio language models.
Technical Environment
  • Core: PyTorch, CUDA, torchaudio/librosa, Hugging Face Transformers.
  • LLM Stack: Large language model backbones, lightweight adapters (LoRA, Q‑Former), instruction tuning pipelines.
  • Audio Models: Neural audio codecs, speech encoders, vocoders, discrete speech tokenizers.
  • Infrastructure: Modern GPU clusters, experiment tracking (Weights & Biases), distributed training frameworks.
  • Deployment: FastAPI/gRPC for services, ONNX/TensorRT for optimized inference.
What We Offer
  • Competitive compensation package with comprehensive benefits.
  • Opportunity to work on cutting‑edge Large Audio Language Models and audio reasoning research with real‑world impact.
  • Collaboration with experienced applied scientists and engineers in speech and multimodal AI.
  • Support for publications at top‑tier conferences and professional development.
  • Access to state‑of‑the‑art GPU infrastructure for training large‑scale audio models.
  • Flexible work arrangements with hybrid/remote options.

Location: Redmond, WA / Palo Alto, CA / Remote

Salary: $150‑$160k Annually.

Centific is an equal‑opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, national origin, ancestry, citizenship status, age, mental or physical disability, medical condition, sex (including pregnancy), gender identity or expression, sexual orientation, marital status, familial status, veteran status or any other characteristic protected by applicable law. We consider qualified applicants regardless of criminal histories, consistent with legal requirements.

#J-18808-Ljbffr