1

Video Captioning Jobs (NOW HIRING)

Vigilantly protect our air by simulating the experience of a viewer and constantly checking for proper audio, video, captioning and picture quality * With the guidance of Technical Managers and ...

Experience with AI-first video workflows (e.g., voice generation, script-to-video, captioning tools) * Production experience: cameras, lighting, directing talent, or overseeing shoots * Experience ...

$84K - $141K/yr

Strong understanding of video production and delivery workflows including formats, frame rates, codecs, compression, color correction, closed captioning, and platform-specific optimization

OBJECTIVES The Video Producer is responsible for creating, capturing, and producing compelling ... Perform color correction, audio enhancement, graphics integration, formatting, captioning, and ...

... captioning etc - be ok with the grunt work amongst the story telling work - Work well with existing ... video content, 30 second commercials, social vids) - High degree of attention to detail ...

OBJECTIVES The Video Producer is responsible for creating, capturing, and producing compelling ... Perform color correction, audio enhancement, graphics integration, formatting, captioning, and ...

Communications Specialist 1

Tualatin, OR ยท On-site

$55K - $73K/yr

Manage video captioning & accessibility Oversee captioning for internal videos to meet accessibility standards. Collaborate with vendors or internal teams for accurate transcription and formatting.

Proven ability to handle the entire lifecycle of a video on your own - concepting, shooting, audio/video editing, captioning, and publishing. * Production & Editing: Comfortable planning and ...

Manage post-production workflows including editing, sound mixing, color correction, captioning and ... Ensure all video content aligns with Pew Research Center's editorial standards, brandguidelinesand ...

Video Production Coordinator

Yorktown, VA ยท On-site

$59K - $69K/yr

CA-VIDEO SERVICES Opening Date: 07/31/2026 Closing Date: Continuous Description Responsible for ... captioning when appropriate. Prepares and distributes completed media across digital platforms ...

Showing results 41-60

Video Captioning information

See salary details

$25K

$74.6K

$160.5K

How much do video captioning jobs pay per year?

As of Sep 14, 2026, the average yearly pay for video captioning in the United States is $74,626.00, according to ZipRecruiter salary data. Most workers in this role earn between $45,000.00 and $94,500.00 per year, depending on experience, location, and employer.

What is video captioning?

Video captioning is the process of transcribing spoken dialogue and relevant audio information from a video into text, which is then displayed on the screen as captions. This helps make video content accessible to people who are deaf or hard of hearing and can also benefit viewers in noisy environments or those who prefer reading along. Captions can be created manually or generated automatically using speech recognition software, and they often include not just spoken words but also important sounds and speaker identification.

What are the key skills and qualifications needed to thrive in video captioning, and why are they important?

To thrive as a Video Captioning Specialist, you need excellent language proficiency, strong attention to detail, and a good understanding of grammar and punctuation, often supported by experience or training in transcription or captioning. Familiarity with captioning software such as Amara, Subtitle Edit, or Aegisub, as well as knowledge of captioning standards and accessibility guidelines, is typically required. Strong time management, adaptability, and communication skills help you meet deadlines and collaborate effectively with content creators. These skills ensure captions are accurate, accessible, and delivered efficiently, which is crucial for audience comprehension and legal compliance.

What are some typical challenges faced by professionals in video captioning, and how can they be overcome?

Professionals in video captioning often encounter challenges such as tight deadlines, ensuring accuracy with fast-paced dialogue, and maintaining consistency with specialized terminology or accents. Overcoming these challenges typically involves using advanced transcription tools, collaborating closely with content creators for clarifications, and maintaining a thorough style guide. Regularly reviewing and updating captioning software skills can also improve efficiency and accuracy, making the workflow smoother and more manageable.

What is the difference between Video Captioning vs Video Transcription?

AspectVideo CaptioningVideo Transcription
CredentialsTypically requires basic language skills, sometimes certification in captioning toolsRequires strong language proficiency, often transcription certifications
Work EnvironmentVideo editing or captioning software, often remoteAudio/video playback, transcription software, remote or office
Industry UsageMedia, entertainment, education, accessibility servicesMedia, legal, medical, general content transcription

Video captioning involves creating timed text overlays for videos to improve accessibility, often requiring familiarity with captioning standards. Video transcription converts spoken content into written text, focusing on accuracy of dialogue or narration. While both roles involve working with audio/video content, captioning emphasizes timing and formatting for viewers, whereas transcription emphasizes verbatim text conversion. Both jobs share skills in language proficiency and often use similar tools, but serve different purposes in media production and accessibility.

More about Video Captioning jobs

What cities are hiring for Video Captioning jobs?

Cities with the most Video Captioning job openings:

What are the most commonly searched types of Video Captioning jobs?

The most popular types of Video Captioning jobs are:

What states have the most Video Captioning jobs?

States with the most job openings for Video Captioning jobs include:

Infographic showing various Video Captioning job openings in the United States as of September 2026, with employment types broken down into 67% Full Time, and 33% Part Time. Highlights an 100% In-person job distribution, with an average salary of $74,626 per year, or $35.9 per hour.

Staff AI Engineer, Applied AI - Smart Vision

Milpitas, CA โ€ข On-site

Arlo Technologies, Inc.
Computer and Electronic Product Manufacturingย โ€ขย 201 - 500 employees

Full-time

Posted 3 days ago

New


Job description

About Arlo:
At Arlo, we're passionate about creating innovative and reliable solutions that help people protect what matters most to them. Our team is dedicated to delivering products that exceed our customers' expectations, while always pushing the boundaries of what's possible in the world of protection technology. We believe that everyone deserves to feel safe and secure, whether they're at home or away, and we're committed to providing our customers with the peace of mind they need to live their lives without worry. Arlo's deep expertise in AI- and CV-powered analytics, cloud services, user experience, product design, and innovative wireless and RF connectivity enables the delivery of a seamless, smart security experience for Arlo users that is easy to set up and interact with every day.

Smart Vision is the AI team behind Arlo's intelligence layer: object and person detection, animal/vehicle/package recognition, custom-trained detections, video captioning and scene description, and natural-language search over a user's video library. Our models run across the edge, the cloud, and third-party foundation models, and they process events from millions of cameras every day.
About the role

As a Staff AI Engineer for Applied AI, you'll be a technical owner of the models behind Arlo's smart features - from computer vision detectors running on-camera to vision-language models that describe what happened, to the retrieval and agent layers that let customers ask questions about their video. You'll pick the right approach for each problem (train, fine-tune, prompt, or retrieve), prove it with solid evals, and take it all the way to production at consumer scale. This is a hands-on applied role: you ship models, not papers.
What you'll do
  • Build, train, and fine-tune computer vision models for detection, classification, tracking, re-identification, and video understanding, and improve them against real-world customer footage - night, weather, motion blur, odd camera angles, edge compute limits.
  • Own our video-understanding pipeline built on vision-language models: frame selection and temporal context, prompt and output-schema design, grounding and hallucination control, multi-event reasoning, and quality tuning for captioning and scene description.
  • Adapt models to our domain: SFT, LoRA/QLoRA, preference tuning, distillation into small deployable models, and knowing when a 200M-parameter specialist beats a frontier model.
  • Own the data and evaluation loop - dataset curation, labeling strategy, hard-negative and failure mining, active learning, benchmark suites, and offline/online metrics that reliably predict customer-perceived quality.
  • Own the embedding and retrieval stack behind video search: multimodal/video embeddings, vector index design and tuning, hybrid search and re-ranking, and natural-language queries over a user's library.
  • Build agentic experiences on top of the vision stack: tool/function calling, multi-step reasoning over event history, RAG and memory, guardrails, and tracing/observability for agent runs.
  • Keep production inference fast and economical - serving stack choice and tuning (vLLM/TensorRT-LLM/Triton), quantization, batching, GPU utilization, and routing between hosted foundation models and self-hosted open models against clear cost and latency targets.
  • Partner with product, data, backend, and firmware/edge teams to turn ambiguous product ideas into shipped AI features, with safe rollout (canary, A/B, feature flags) and real production telemetry.
  • Raise the engineering bar - architecture and design reviews, MLOps practices, documentation, and mentoring senior and mid-level engineers.

What we're looking for
  • BS in Computer Science or a related technical field with 8+ years of experience; MS/PhD in ML, CV, or a related field preferred (or equivalent practical experience).
  • 8+ years building production ML/AI systems, with a track record of owning models end to end - problem framing, data, training, evaluation, deployment, iteration.
  • Strong computer vision depth: detection, classification, segmentation, tracking, video understanding; you've trained and shipped CV models in a real product.
  • 3+ years working with LLMs or VLMs in production - multimodal modeling, prompt and context design, fine-tuning, and evaluation.
  • Strong Python + PyTorch; solid distributed systems and cloud fundamentals (AWS, Docker).
  • Rigorous about evaluation and data quality - you build the benchmark before you build the model.
  • Experience with embeddings and vector search at scale; multimodal or video retrieval strongly preferred.
  • Experience shipping LLM agents with tool use, and pragmatic judgment about when not to use an agent.
  • Working knowledge of inference optimization (serving stacks, quantization, batching, GPU performance) and cost/latency tradeoffs at scale.
  • Demonstrated technical leadership without formal authority: influencing roadmaps, mentoring engineers, driving cross-team decisions.
  • Bias to ship, comfort with ambiguity, strong written communication.

Nice to have
  • Edge/on-device inference (quantization-aware training, pruning, NPU/DSP toolchains) and streaming inference.
  • Video processing at scale (decoding, frame sampling, FFmpeg-class tooling).
  • Audio or sensor-fusion models complementing video; privacy-preserving or on-device personalization.
  • Open-source contributions to CV, serving, retrieval, or agent frameworks; consumer IoT or camera/security domain experience.

We're committed to inclusivity and selecting the strongest candidate-no matter their background. Even if you don't meet every listed qualification, we encourage you to apply. We're happy to support growth in areas essential to the role. Interested in learning more about our workplace? Visit and follow our LinkedIn, and Glassdoor pages to read employee insights and get updates of what it's like to be part of Arlo.

Arlo is proud to be an Equal Opportunity Employer. We value inclusion and are committed to inclusive, and harassment-free workplace. We prohibit discrimination and harassment based on all legally protected statuses in all hiring and employment.

We provide reasonable accommodations to applicants and employees with disabilities, who are pregnant or have a related medical condition, or who have sincerely held religious beliefs, observances, and practices. Pursuant to applicable state and municipal Fair Chance Laws and Ordinances, the Company will consider for employment qualified applicants with arrest and conviction records.