1

Video Captioning Jobs (NOW HIRING)

Captioning Technician Job Category: Student Hourly Job Profile: Student Help Job Summary ... Compensation $12 per hour Hours 10-25 per week This position ensures video captions in Kaltura are ...

This involves ftp ingestion, quality checks, captioning, editing, and DVD authoring. Key ... Author and edit video and iconographic films to wrap up productions as needed * Manage the ...

Create closed-captioning files and load final files into the server for publishing. Understand, anticipate, and define plan for executing video post-production to showcase features and benefits of ...

... audio, video, captioning and picture quality With the guidance of Technical Managers and supervisors, work as a team to promptly recover from any on-air discrepancies Other duties as assigned ...

next page

Showing results 1-20

Video Captioning information

See salary details

$25K

$74.6K

$160.5K

How much do video captioning jobs pay per year?

As of Aug 4, 2026, the average yearly pay for video captioning in the United States is $74,626.00, according to ZipRecruiter salary data. Most workers in this role earn between $45,000.00 and $94,500.00 per year, depending on experience, location, and employer.

What are some typical challenges faced by professionals in video captioning, and how can they be overcome?

Professionals in video captioning often encounter challenges such as tight deadlines, ensuring accuracy with fast-paced dialogue, and maintaining consistency with specialized terminology or accents. Overcoming these challenges typically involves using advanced transcription tools, collaborating closely with content creators for clarifications, and maintaining a thorough style guide. Regularly reviewing and updating captioning software skills can also improve efficiency and accuracy, making the workflow smoother and more manageable.

Can I get paid to caption videos?

Video captioning is a legitimate job that involves creating accurate subtitles for videos, often requiring skills in transcription and familiarity with captioning tools. Many companies and freelance platforms offer paid opportunities for captioners, with pay rates varying based on experience and project complexity.

What is the difference between Video Captioning vs Video Transcription?

AspectVideo CaptioningVideo Transcription
CredentialsTypically requires basic language skills, sometimes certification in captioning toolsRequires strong language proficiency, often transcription certifications
Work EnvironmentVideo editing or captioning software, often remoteAudio/video playback, transcription software, remote or office
Industry UsageMedia, entertainment, education, accessibility servicesMedia, legal, medical, general content transcription

Video captioning involves creating timed text overlays for videos to improve accessibility, often requiring familiarity with captioning standards. Video transcription converts spoken content into written text, focusing on accuracy of dialogue or narration. While both roles involve working with audio/video content, captioning emphasizes timing and formatting for viewers, whereas transcription emphasizes verbatim text conversion. Both jobs share skills in language proficiency and often use similar tools, but serve different purposes in media production and accessibility.

What are the key skills and qualifications needed to thrive in video captioning, and why are they important?

To thrive as a Video Captioning Specialist, you need excellent language proficiency, strong attention to detail, and a good understanding of grammar and punctuation, often supported by experience or training in transcription or captioning. Familiarity with captioning software such as Amara, Subtitle Edit, or Aegisub, as well as knowledge of captioning standards and accessibility guidelines, is typically required. Strong time management, adaptability, and communication skills help you meet deadlines and collaborate effectively with content creators. These skills ensure captions are accurate, accessible, and delivered efficiently, which is crucial for audience comprehension and legal compliance.

What is video captioning?

Video captioning is the process of transcribing spoken dialogue and relevant audio information from a video into text, which is then displayed on the screen as captions. This helps make video content accessible to people who are deaf or hard of hearing and can also benefit viewers in noisy environments or those who prefer reading along. Captions can be created manually or generated automatically using speech recognition software, and they often include not just spoken words but also important sounds and speaker identification.
More about Video Captioning jobs
What cities are hiring for Video Captioning jobs? Cities with the most Video Captioning job openings:
What are the most commonly searched types of Video Captioning jobs? The most popular types of Video Captioning jobs are:
What states have the most Video Captioning jobs? States with the most job openings for Video Captioning jobs include:
Infographic showing various Video Captioning job openings in the United States as of July 2026, with employment types broken down into 23% Internship, 37% As Needed, 24% Full Time, 5% Part Time, 1% Contract, and 10% Nights. Highlights an 91% Physical, 3% Hybrid, and 6% Remote job distribution, with an average salary of $74,626 per year, or $35.9 per hour.

Machine Learning Engineer (Video Understanding & Segmentation)

Maxinsights Corporation

Santa Clara, CA • On-site

Full-time

Medical, Dental, Vision, Retirement, PTO

Posted 20 days ago


Job description

Job Description:

We are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale — turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.

Responsibilities

  • Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.

  • Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.

  • Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.

  • Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.

  • Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.

  • Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.

  • Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.

  • Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.

  • Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.

  • Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.

Minimum Qualifications

  • MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.

  • 3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.

  • Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.

  • Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.

  • Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).

  • Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).

  • Experience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).

Preferred Qualifications

  • PhD with a research focus in video understanding, multi-modal learning, or vision-language models.

  • Experience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.

  • Experience working with egocentric video datasets or head-mounted-device (HMD) captured data.

  • Track record of deploying production-scale video search or retrieval systems.

  • Experience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.

  • Publications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).

  • Experience with humanoid robotics or embodied AI data pipelines is a plus.

Default Benefits:
  • Health insurance

  • Vision care

  • Dental coverage:

  • 401(k)

  • Paid holidays

  • PTO (Paid Time Off)

  • Sick leave