1

Audio Transcription Jobs (NOW HIRING)

Be Seen First

This is not a general audio-production or speech-transcription role. Candidates should have strong musical listening, notation, and technical problem-solving abilities. What you'll work on:

Showing results 21-40

Audio Transcription information

See salary details

$17

$22

$29

How much do audio transcription jobs pay per hour?

As of Sep 5, 2026, the average hourly pay for audio transcription in the United States is $22.56, according to ZipRecruiter salary data. Most workers in this role earn between $21.15 and $23.80 per hour, depending on experience, location, and employer.

What is audio transcription?

Audio transcription is the process of converting spoken words from audio recordings into written text. This service is often used for interviews, podcasts, meetings, legal proceedings, and other situations where a written record of spoken content is needed. Transcription can be done manually by a person or automatically using speech recognition software, although human transcription is usually more accurate. Accurate audio transcription requires good listening skills, attention to detail, and sometimes knowledge of specialized terminology depending on the subject matter.

What is an audio transcriptionist?

The duties of an audio transcription professional involve working to transcribe speech to text. They listen to a recording or audio files or type what a person is saying in real time, focusing on creating accurate text of audio. Transcriptionists may work in the healthcare industry, in a courtroom, or a corporate board room. They make transcripts of speeches, news reports on TV or radio, or conference calls. An audio transcription professional may also work in law enforcement where they record police radio interactions. The qualifications that you need to work in audio transcription vary.

What are some common challenges faced by audio transcriptionists, and how can they be managed?

Audio transcriptionists often encounter challenges such as poor audio quality, heavy accents, overlapping speakers, and specialized terminology. Managing these difficulties typically involves using high-quality headphones, leveraging transcription software with playback controls, and frequently pausing or rewinding to capture unclear sections. Additionally, building a glossary of industry-specific terms and collaborating with clients or colleagues for clarifications can significantly improve accuracy and efficiency.

How do I start a job transcription?

To start a job in audio transcription, you should develop strong listening and typing skills, familiarize yourself with transcription software and tools, and create a professional resume highlighting relevant experience. Many transcription jobs require passing a skills test or sample transcription to demonstrate accuracy and speed before being hired.

What cities are hiring for Audio Transcription jobs?

Cities with the most Audio Transcription job openings:

What are the most commonly searched types of Audio Transcription jobs?

The most popular types of Audio Transcription jobs are:

What states have the most Audio Transcription jobs?

States with the most job openings for Audio Transcription jobs include:

What job categories do people searching Audio Transcription jobs look for?

The top searched job categories for Audio Transcription jobs are:

Infographic showing various Audio Transcription job openings in the United States as of August 2026, with employment types broken down into 86% Full Time, and 14% Part Time. Highlights an 72% In-person, 14% Hybrid, and 14% Remote job distribution, with an average salary of $46,930 per year, or $22.6 per hour.

Senior Software Engineer - AI Hardware Backend

Nuha

Manhattan, NY โ€ข On-site

$120 - $160/hr

Other

Re-posted yesterday


Key responsibilities

  • Build and maintain FastAPI services that manage WebSocket connections, device sessions, and conversation state

  • Orchestrate LLMs, process sensor data, and connect embedded devices to AI infrastructure

  • Develop database architecture for user profiles, interaction history, conversation memories, and sensor logs


Job description

Senior Software Engineer - AI Hardware Backend

Location: NYC - In office.

About Us

Nuha is a research lab based in New York City. We teach companions how to understand, interact with, and respond to humans and the physical world.

Your Role

You'll own the backend infrastructure that powers our AI hardware. This means building the FastAPI services that orchestrate LLMs, audio transcription, sensor data processing, and model inference, connecting our embedded devices to the intelligence that makes them work.

You'll work directly with our hardware team and our data/R&D team to turn sensor streams and voice interactions into intelligent responses.

What You'll Build
  • FastAPI services managing WebSocket connections, device sessions, and conversation state
  • LLM orchestration (Groq, OpenAI, Modal), conversation memory, prompt engineering, safety guards
  • Database architecture (Supabase) - user profiles, interaction history, conversation memories, sensor logs
  • Authentication, connection lifecycle, monitoring, and error recovery across production devices
  • Process streaming mmWave radar frames and generate tracking visualizations (VisPy)
  • Fuse audio, radar, and accelerometer data into coherent context for LLMs
  • Build sensor summarization pipelines using VLMs to describe physical space
  • Event-driven coordination of audio capture, VAD, TTS playback, and sensor streams
Product Features & Infrastructure
  • Ship new interaction features based on customer feedback and product vision
  • Build internal tools for debugging device interactions and analyzing sensor data in production
  • Design data pipelines for user analytics, interaction logging, and model evaluation
  • Performance optimization - reduce latency, improve reliability, handle connection failures gracefully
  • Testing infrastructure for hardware-in-the-loop validation and API integration tests
  • Iterate on the interaction experience - make conversations feel more natural and responsive
Requirements
  • Systems thinking - Bridge hardware constraints, API limitations, and ML requirements
  • Real-time data pipelines - Experience with streaming data, low-latency systems, or sensor processing
  • API orchestration - Coordinate multiple services (transcription, LLM, inference) reliably
  • Cross-functional collaboration - Work with hardware engineers and ML researchers effectively
  • Ships production code - Build reliable services that handle real user interactions
Bonus
  • Mathematical maturity
  • VisPy/graphics experience
  • Hardware experience
Team

You'll work alongside engineers doing: RF/mmWave radar design, PCB/enclosure design, embedded client software, ML accelerator R&D, data infrastructure, and manufacturing systems.

Interview Process
  • Phone screen
  • Technical interview - Standard SWE problem solving and system design
  • In-person deep dive - Walk us through a complex technical problem you've solved
  • Work trial - Spend time with the team working on a real problem
Location

NYC - In office. We're building a physical product and need to work together in person.

  • Any backend/systems projects you've built
#J-18808-Ljbffr