1

Audio Annotation Jobs (NOW HIRING)

... and annotation * Set new standards for how we evaluate and benchmark our audio models What You ... Bring * Strong applied mindset and ability to balance scientific novelty with product impact.

Work with technical staff to improve annotation tools for efficient audio workflows. BASIC QUALIFICATIONS: * Native proficiency in Chinese with exposure to diverse accents, dialects, or regional ...

Work with technical staff to improve annotation tools for efficient audio workflows. BASIC QUALIFICATIONS: * Native proficiency in Arabic with exposure to diverse accents, dialects, or regional ...

Work with technical staff to improve annotation tools for efficient audio workflows. BASIC QUALIFICATIONS: * Native proficiency in Punjabi with exposure to diverse accents, dialects, or regional ...

Work with technical staff to improve annotation tools for efficient audio workflows. BASIC QUALIFICATIONS: * Native proficiency in Portuguese with exposure to diverse accents, dialects, or regional ...

Work with technical staff to improve annotation tools for efficient audio workflows. BASIC QUALIFICATIONS: * Native proficiency in Indonesian with exposure to diverse accents, dialects, or regional ...

Work with technical staff to improve annotation tools for efficient audio workflows. BASIC QUALIFICATIONS: * Native proficiency in Polish with exposure to diverse accents, dialects, or regional ...

Work with technical staff to improve annotation tools for efficient audio workflows. BASIC QUALIFICATIONS: * Native proficiency in Danish with exposure to diverse accents, dialects, or regional ...

Technical Program Manager, Data

San Francisco, CA · On-site

$152K - $196K/yr

Responsibilities : • Lead audio data collection and annotation efforts at Sesame. • Collaborate with research and product teams to understand and formalize their requirements. • Identify and ...

Lead audio data collection and annotation efforts at Sesame. * Collaborate with research and product teams to understand and formalize their requirements. * Identify and manage internal resources and ...

Data Ops Lead

New York, NY · On-site +1

$150K - $190K/yr

About the role Your mission is to turn Neon's raw consumer audio streams into the cleanest, most ... Standing up human transcription, annotation and other operations, largely overseas, that make it ...

Showing results 21-40

Audio Annotation information

See salary details

$29.5K

$84.5K

$171.5K

How much do audio annotation jobs pay per year?

As of Aug 8, 2026, the average yearly pay for audio annotation in the United States is $84,456.00, according to ZipRecruiter salary data. Most workers in this role earn between $50,000.00 and $113,000.00 per year, depending on experience, location, and employer.

What are the key skills and qualifications needed to thrive as an audio annotator, and why are they important?

To thrive as an Audio Annotator, you need strong attention to detail, excellent listening skills, and familiarity with linguistic concepts, often supported by relevant coursework or experience in linguistics or audio processing. Proficiency in annotation tools such as ELAN, Audacity, or Praat, as well as experience with data labeling platforms, is typically required. Strong organizational skills, patience, and the ability to work independently make someone stand out in this role. These skills ensure accurate and consistent audio data labeling, which is essential for training reliable AI and speech recognition systems.

What are some common challenges faced by audio annotators, and how can they be managed effectively?

Audio annotators often encounter challenges such as distinguishing overlapping voices, dealing with low-quality recordings, and maintaining consistency in labeling. To manage these, it's important to use high-quality headphones, familiarize yourself with annotation guidelines, and communicate regularly with your team to resolve ambiguities. Many organizations also provide regular feedback sessions and quality checks to ensure accuracy and support continuous improvement.

What is audio annotation?

Audio annotation is the process of labeling or tagging audio data with relevant information, such as identifying sounds, speech, speakers, or background noises. This process helps train machine learning models to recognize and understand audio content. Audio annotation can involve tasks like transcribing speech, marking segments with specific sounds, or categorizing audio clips by genre or emotion. It is widely used in developing applications for speech recognition, virtual assistants, and audio analysis.
More about Audio Annotation jobs
What cities are hiring for Audio Annotation jobs? Cities with the most Audio Annotation job openings:
What states have the most Audio Annotation jobs? States with the most job openings for Audio Annotation jobs include:
What job categories do people searching Audio Annotation jobs look for? The top searched job categories for Audio Annotation jobs are:
Infographic showing various Audio Annotation job openings in the United States as of August 2026, with employment types broken down into 33% Full Time, 33% Part Time, and 34% Contract. Highlights an 33% In-person, and 67% Remote job distribution, with an average salary of $84,456 per year, or $40.6 per hour.

Applied Researcher, Audio

Cartesia

San Francisco, CA • On-site

$200K - $350K/yr

Full-time

Medical, Dental, Vision, Retirement, PTO

Re-posted 24 days ago


Job description

About Cartesia
Our mission is to architect AI that learns from and interacts with the world like humans do.
We're pioneering the model architectures that will make this possible. Our founding team met as PhDs at the Stanford AI Lab, where we invented State Space Models or SSMs, a new primitive for training efficient, large-scale foundation models. Our team combines deep expertise in model innovation and systems engineering paired with a design-minded product engineering team to build and ship cutting edge models and experiences.
We're funded by leading investors at Index Ventures and Lightspeed Venture Partners, along with Factory, Conviction, A Star, General Catalyst, SV Angel, Databricks and others. We're fortunate to have the support of many amazing advisors, and 90+ angels across many industries, including the world's foremost experts in AI.
About the Role
You will be responsible for leading research at the frontier of realtime conversation and human AI interaction. You will contribute across the stack to novel architectures, data, and evals for realtime audio, and translate them into state-of-the-art models used in voice agents around the world.
Your Impact
  • Architect and develop new architectures for realtime audio understanding, generation, and speech-to-speech models that reason jointly over multiple modalities in realtime
  • Contribute to frontier multimodal and multilingual datasets for pre-training and post-training, including curating data mixes and developing new methods for synthetic data generation and annotation
  • Set new standards for how we evaluate and benchmark our audio models

What You Bring
  • Strong applied mindset and ability to balance scientific novelty with product impact.
  • Excited and able to work across the stack from infra, to data, to evals, to architecture to solve customer problems and build state-of-the-art models.
  • Deep expertise in deep generative modeling. Previous experience in audio understanding, audio generation, speech-to-speech, or language modeling preferred but not required.
  • Experience with large-scale training, GPU/TPU acceleration, and model optimization.

Note: Cartesia participates in E-Verify and will provide the federal government with Form I-9 information to confirm employment eligibility after hire.
More Details
In-office policy: We're an in-person team based out of offices in S San Francisco, GB London and I Bangalore. We love being in the office, hanging out together, and learning from each other every day.
Visa sponsorship: We provide visa sponsorship support and assess each circumstance on a case-by-case basis. However, visa sponsorship is dependent on many factors, including the role you are applying for, and the location you are going to be based, and so we can't always guarantee success. Your Recruiter will work with you to understand your visa sponsorship needs from the first call.
We ship fast. All of our work is novel and cutting edge, and execution speed is paramount. We have a high bar, and we don't sacrifice quality or design along the way.
We support each other. We have an open & inclusive culture that's focused on giving everyone the resources they need to succeed.
Our Benefits (US Employees Only)
Compensation Competitive base salary alongside attractive equity package.
Health Insurance Fully covered medical insurance along with dental and vision for you and your family.
Parental Leave 9 weeks paternity & 12 weeks maternity leave
401(k)
Commuter Allowance A monthly stipend to help you get to and from the office.
Flexible PTO Take as much time as you need to recharge your batteries.
Meals & Snacks Lunch, dinner and plenty of snacks, provided daily.
Your own personal Yoshi
Our Commitment to Equal Opportunity
Cartesia is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, national origin, age, disability, veteran status, genetic information, or any other legally protected status.