1

Ai Audio Jobs in Virginia (NOW HIRING)

Audio AI Engineer, #1085 Multilingual Speech-to-Text Engineer - On-Device Model Optimization, #1085 A Role with Purpose and Impact This role builds the speech recognition core of a mobile translation ...

Audio AI Engineer, #1085 Multilingual Speech-to-Text Engineer -- On-Device Model Optimization, #1085 A Role with Purpose and Impact This role builds the speech recognition core of a mobile ...

Audio AI Engineer

Reston, VA · On-site

$80K - $160K/yr

... Audio AI Engineer, #1085 Multilingual Speech-to-Text Engineer - On-Device Model Optimization, #1085 A Role with Purpose and Impact This role builds the speech recognition core of a mobile translation ...

AI/ML Engineer

Reston, VA · On-site

$175K - $220K/yr

Our model enables our customers to perform critical data labeling tasks over multi-media sources such as video, images, audio, and documents. Our AI/ML Engineer is the core mission specialist who ...

As AI becomes central to critical infrastructure, advancing its security and resilience offers a ... audio analysis) * Experience and knowledge in cybersecurity best practices * Demonstrated ability ...

next page

Showing results 1-20

Ai Audio information

How does an AI audio specialist typically collaborate with other teams during a project?

AI Audio specialists often work closely with product managers, software engineers, and UX designers to integrate audio solutions into products or platforms. Collaboration usually involves joint meetings to define project requirements, frequent communication to troubleshoot integration issues, and sharing feedback to refine audio models. A successful AI Audio professional is proactive in bridging technical and creative perspectives, ensuring that audio features meet user needs and technical standards. This team-oriented environment fosters learning and can open pathways to roles in project leadership or advanced technical development.

What is the difference between Ai Audio vs Voiceover Artist?

AspectAi AudioVoiceover Artist
Required CredentialsTechnical skills, AI and audio editing knowledgeVoice training, acting skills, demo reels
Work EnvironmentDigital, remote, tech-focusedRecording studios, remote, live performances
Industry UsageMedia production, AI development, tech companiesAdvertising, entertainment, media
Search & Comparison IntentTechnical, AI-driven audio solutionsCreative voice work, acting

Ai Audio involves creating audio content using artificial intelligence technology, focusing on automation and digital tools. Voiceover Artists provide human voice recordings for various media, emphasizing performance and acting skills. While Ai Audio is tech-based and automated, Voiceover Artists rely on vocal talent and creativity. Both roles are essential in media production but serve different purposes and skill sets.

What are the key skills and qualifications needed to thrive as an AI audio engineer?

To thrive as an AI Audio Engineer, you need strong foundations in audio signal processing, machine learning, and computer programming, often supported by a degree in computer science, electrical engineering, or a related field. Familiarity with tools such as Python, TensorFlow or PyTorch, digital audio workstations (DAWs), and version control systems is typically required. Attention to detail, creativity, and effective collaboration are standout soft skills for this role. These skills and qualities are essential for developing innovative audio solutions and ensuring seamless integration of AI technologies in audio applications.

What is an AI audio engineer?

An AI Audio Engineer is a professional who uses artificial intelligence and machine learning technologies to develop, enhance, or analyze audio content. Their work may include creating AI-driven tools for music production, speech recognition, audio restoration, or sound synthesis. AI Audio Engineers often collaborate with software developers, musicians, and sound designers to integrate intelligent audio solutions into various applications. This field requires knowledge of audio engineering, signal processing, and programming skills.
What cities in Virginia are hiring for Ai Audio jobs? Cities in Virginia with the most Ai Audio job openings:
Infographic showing various Ai Audio job openings in Virginia as of August 2026, with employment types broken down into 79% Full Time, 18% Part Time, and 3% Contract. Highlights an 67% Physical, 4% Hybrid, and 29% Remote job distribution.

$80K - $160K/yr

Full-time

Posted 5 days ago


Job description

Audio AI Engineer, #1085

Multilingual Speech-to-Text Engineer - On-Device Model Optimization, #1085

A Role with Purpose and Impact

This role builds the speech recognition core of a mobile translation capability supporting a government agency's national security mission. The engineer will take large, high-quality speech-to-text models spanning many language families and adapt, compress, and optimize them so they run performantly on an iPhone - including handling the reality that speakers frequently mix in borrowed English terms mid-utterance, and the model needs to make a sound call on whether to transcribe those terms in English or in the source language's own transliteration.

This is an applied ML role, not a research-only position. The strongest candidate can move fluidly from raw audio data, to model adaptation and compression experiments, to a rigorous evaluation framework - and can clearly explain what they're building, why it's better than the status quo, and how they'll know it worked.

What This Role Is (and Isn't)

This position owns the speech-to-text model - its data, its training/adaptation, its size and latency on-device, and its accuracy across languages. It does not own iOS application development, translation (source-language-to-target-language), or the Swift/AVFoundation integration layer; those are handled by a separate mobile engineering function this role will collaborate closely with.

Key Responsibilities

  • Data pipelines: Ingest, clean, segment, label, and version multilingual audio and transcript data, with attention to code-switching and borrowed-word phenomena across the target language set.
  • Model adaptation: Fine-tune and compress large ASR models (using LoRA/QLoRA, quantization, distillation, or other parameter-efficient and size-reduction techniques as appropriate) to fit iPhone-class memory, latency, and battery constraints, while preserving transcription quality.
  • Dynamic, per-language deployment: Design model packaging so language-specific weights can be selected and downloaded on demand based on use-case context (e.g., an operator interviewing a Chinese speaker pulls only the Chinese ASR weights).
  • Loanword/transliteration handling: Build and evaluate model behavior for deciding when a borrowed English term should be transcribed as-is versus rendered in the source language's transliteration or native equivalent.
  • Evaluation: Build reproducible evaluation pipelines (word/character error rate, latency, robustness to accent/noise/speaking rate/code-switching) and clearly articulate results against defined success criteria for each language and deployment target.
  • Documentation & communication: Produce clear model cards, dataset documentation, and evaluation write-ups that let technical and non-technical stakeholders understand what the model does, how it compares to alternatives, and what its risks and limitations are.

Required Qualifications

  • Bachelor's degree in Computer Science, Data Science, Machine Learning, Computational Linguistics, or a closely related field.
  • Strong data-engineering background building production pipelines for large, messy, or unstructured audio/text datasets.
  • Hands-on experience fine-tuning or adapting speech/audio models using parameter-efficient methods (LoRA, QLoRA, adapters) and/or model compression techniques (quantization, distillation, pruning) for constrained hardware.
  • Practical experience with ASR/speech-to-text model development and evaluation across multiple languages, including error analysis under real-world conditions (accents, noise, code-switching).
  • Strong Python and SQL skills; experience with PyTorch, Hugging Face Transformers/PEFT, torchaudio, librosa, or comparable tooling.
  • Experience deploying and monitoring production ML systems, with an understanding of secure handling of sensitive audio, transcripts, and derived data in a regulated environment.
  • Ability to clearly explain model behavior, tradeoffs, and limitations to both technical and non-technical stakeholders.

Preferred (Not Required)

  • Prior exposure to mobile/on-device ML deployment constraints (even without owning the mobile codebase directly).
  • Experience with agentic or multi-step workflow orchestration involving model outputs, retrieval, or human review.

The estimated salary range for this position is $80,000 - $160,000. This salary range is not a guarantee of compensation. The offered salary will be based on factors including relevant experience, geographic location, internal equity, and applicable contractual requirements. *Compensation may fall outside this range when appropriate.