Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes. * Architect agentic systems and ...
Quick apply
Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes. * Architect agentic systems and ...
Quick apply
Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes. * Architect agentic systems and ...
Orem, UT · On-site
$13 - $21.65/hr
Familiarity with document authoring tools and video captioning tools (training can be provided). * Sensitivity to the needs of individuals with disabilities. Knowledge / Skills / Abilities * Ability ...
Orem, UT · On-site
$13 - $21.65/hr
Familiarity with document authoring tools and video captioning tools (training can be provided). * Sensitivity to the needs of individuals with disabilities. Knowledge / Skills / Abilities * Ability ...
Stevens Point, WI · On-site
$12/hr
Captioning Technician Job Category: Student Hourly Job Profile: Student Help Job Summary ... Compensation $12 per hour Hours 10-25 per week This position ensures video captions in Kaltura are ...
Stevens Point, WI · On-site
$12/hr
Captioning Technician Job Category: Student Hourly Job Profile: Student Help Job Summary ... Compensation $12 per hour Hours 10-25 per week This position ensures video captions in Kaltura are ...
New York, NY · Remote
$113K - $132K/yr
Dotsub is building a unique video captioning and translation management tool that enables clients organize their in-house linguists, create and manage custom workflows, take advantage of the power of ...
New York, NY · Remote
$113K - $132K/yr
Dotsub is building a unique video captioning and translation management tool that enables clients organize their in-house linguists, create and manage custom workflows, take advantage of the power of ...
Captioning Technician Job Category: Student Hourly Job Profile: Student Help Job Summary ... Compensation $12 per hour Hours 10-25 per week This position ensures video captions in Kaltura are ...
Captioning Technician Job Category: Student Hourly Job Profile: Student Help Job Summary ... Compensation $12 per hour Hours 10-25 per week This position ensures video captions in Kaltura are ...
New York, NY · On-site +1
$113K - $132K/yr
Dotsub is building a unique video captioning and translation management tool that enables clients organize their in-house linguists, create and manage custom workflows, take advantage of the power of ...
New York, NY · On-site +1
$113K - $132K/yr
Dotsub is building a unique video captioning and translation management tool that enables clients organize their in-house linguists, create and manage custom workflows, take advantage of the power of ...
Holland, OH · On-site
This involves ftp ingestion, quality checks, captioning, editing, and DVD authoring. Key ... Author and edit video and iconographic films to wrap up productions as needed * Manage the ...
Holland, OH · On-site
This involves ftp ingestion, quality checks, captioning, editing, and DVD authoring. Key ... Author and edit video and iconographic films to wrap up productions as needed * Manage the ...
... video captioning, speech-to-text generation. Preferred : • Publications in top-tier venues demonstrating your expertise in multimodal AI research. • Experience in writing efficient GPU kernels ...
... video captioning, speech-to-text generation. Preferred : • Publications in top-tier venues demonstrating your expertise in multimodal AI research. • Experience in writing efficient GPU kernels ...
Sterling, VA · On-site +1
$40 - $70/hr
Required Skills * 5+ years of experience in accessibility testing, QA, compliance testing, or video/TV device testing * Hands-on experience evaluating closed captioning quality * Strong understanding ...
Sterling, VA · On-site +1
$40 - $70/hr
Required Skills * 5+ years of experience in accessibility testing, QA, compliance testing, or video/TV device testing * Hands-on experience evaluating closed captioning quality * Strong understanding ...
Required Skills * 5+ years of experience in accessibility testing, QA, compliance testing, or video/TV device testing * Hands-on experience evaluating closed captioning quality * Strong understanding ...
Quick apply
Required Skills * 5+ years of experience in accessibility testing, QA, compliance testing, or video/TV device testing * Hands-on experience evaluating closed captioning quality * Strong understanding ...
Jr. Video Editor - Sabey Corporation ID 102539 Application Deadline 8/22/2026 Company Sabey ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Jr. Video Editor - Sabey Corporation ID 102539 Application Deadline 8/22/2026 Company Sabey ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Sabey is growing its in-house video team and looking for creative professionals who are passionate ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Sabey is growing its in-house video team and looking for creative professionals who are passionate ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Sabey is growing its in-house video team and looking for creative professionals who are passionate ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Sabey is growing its in-house video team and looking for creative professionals who are passionate ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Troutman, NC · On-site
Create closed-captioning files and load final files into the server for publishing. Understand, anticipate, and define plan for executing video post-production to showcase features and benefits of ...
Troutman, NC · On-site
Create closed-captioning files and load final files into the server for publishing. Understand, anticipate, and define plan for executing video post-production to showcase features and benefits of ...
Sabey is growing its in-house video team and looking for creative professionals who are passionate ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Sabey is growing its in-house video team and looking for creative professionals who are passionate ... Applying color correction, audio mixing, motion graphics, lower thirds, and captioning that align ...
Charlotte, NC · On-site
... audio, video, captioning and picture quality With the guidance of Technical Managers and supervisors, work as a team to promptly recover from any on-air discrepancies Other duties as assigned ...
Charlotte, NC · On-site
... audio, video, captioning and picture quality With the guidance of Technical Managers and supervisors, work as a team to promptly recover from any on-air discrepancies Other duties as assigned ...
Edit short- and long-form video content for digital, social media, websites, presentations, and ... Perform basic color correction, audio cleanup, captioning, and quality control to ensure polished ...
Edit short- and long-form video content for digital, social media, websites, presentations, and ... Perform basic color correction, audio cleanup, captioning, and quality control to ensure polished ...
Edit short- and long-form video content for digital, social media, websites, presentations, and ... Perform basic color correction, audio cleanup, captioning, and quality control to ensure polished ...
Edit short- and long-form video content for digital, social media, websites, presentations, and ... Perform basic color correction, audio cleanup, captioning, and quality control to ensure polished ...
Ensure all video content meets accessibility and ADA compliance standards, including captioning * Manage video equipment, footage organization, file storage, and content archives * Balance multiple ...
Ensure all video content meets accessibility and ADA compliance standards, including captioning * Manage video equipment, footage organization, file storage, and content archives * Balance multiple ...
Burlington, VT · On-site
$24 - $27.50/hr
Knowledge of accessibility best practices for events (e.g., captioning, assistive listening ... Position Information Position Title Video and Broadcasting OC2 N Posting Number S6216PO Department ...
Burlington, VT · On-site
$24 - $27.50/hr
Knowledge of accessibility best practices for events (e.g., captioning, assistive listening ... Position Information Position Title Video and Broadcasting OC2 N Posting Number S6216PO Department ...
$25K - $37.3K
18% of jobs
$46.2K is the 25th percentile. Wages below this are outliers.
$37.3K - $49.6K
10% of jobs
$49.6K - $62K
20% of jobs
The median wage is $63.6K / yr.
$62K - $74.3K
20% of jobs
$74.3K - $86.6K
3% of jobs
$90.2K is the 75th percentile. Wages above this are outliers.
$86.6K - $98.9K
16% of jobs
$98.9K - $111.2K
9% of jobs
$111.2K - $123.5K
0% of jobs
$123.5K - $135.9K
0% of jobs
$135.9K - $148.2K
0% of jobs
$148.2K - $160.5K
4% of jobs
$25K
$74.6K
$160.5K
| Aspect | Video Captioning | Video Transcription |
|---|---|---|
| Credentials | Typically requires basic language skills, sometimes certification in captioning tools | Requires strong language proficiency, often transcription certifications |
| Work Environment | Video editing or captioning software, often remote | Audio/video playback, transcription software, remote or office |
| Industry Usage | Media, entertainment, education, accessibility services | Media, legal, medical, general content transcription |
Video captioning involves creating timed text overlays for videos to improve accessibility, often requiring familiarity with captioning standards. Video transcription converts spoken content into written text, focusing on accuracy of dialogue or narration. While both roles involve working with audio/video content, captioning emphasizes timing and formatting for viewers, whereas transcription emphasizes verbatim text conversion. Both jobs share skills in language proficiency and often use similar tools, but serve different purposes in media production and accessibility.

Santa Clara, CA • On-site
Full-time
Medical, Dental, Vision, Retirement, PTO
Posted 20 days ago
We are seeking a highly motivated Machine Learning Engineer to join our core research and development team, focused on video understanding and segmentation. In this role, you will build the systems that let us search, decompose, and describe massive volumes of egocentric and human-robot video at scale — turning raw, unstructured footage into structured, searchable, and richly annotated training data. You will work across video/image embedding models, LLM-based video understanding, and agentic pipelines that orchestrate multiple models into end-to-end workflows. This is a foundational role that directly shapes the data quality and scalability of our entire training data platform.
Responsibilities
Build and optimize video/image embedding pipelines using CLIP-style and other vision-language embedding models to power large-scale, multi-modal video search and retrieval.
Develop LLM-based video understanding systems for semantic indexing, summarization, and question-answering over long-form egocentric and third-person video.
Design and implement instruction-level and action-level video chunking/segmentation algorithms that decompose long videos into structured, temporally-aligned clips.
Build automated video captioning systems that combine vision-language models and LLMs to produce fine-grained, temporally-grounded descriptions of actions and scenes.
Architect agentic systems and orchestration pipelines that chain embedding, captioning, retrieval, and LLM reasoning steps into reliable, end-to-end video understanding workflows.
Develop and scale video search infrastructure (vector indexing, retrieval, ranking) to support semantic and multi-modal queries over millions of video clips.
Collaborate with annotation, data engineering, and robotics teams to integrate video understanding outputs into downstream training pipelines for embodied AI and robot learning.
Evaluate and benchmark embedding models, LLMs, and agentic frameworks against production needs; track frontier research and bring relevant techniques into the platform.
Contribute to internal tooling, documentation, patents, and open-source initiatives where applicable.
Mentor junior engineers and interns, and help shape the long-term technical roadmap for video understanding.
Minimum Qualifications
MS or PhD in Computer Science, Electrical Engineering, or a related technical field, or equivalent practical experience.
3+ years of hands-on experience in computer vision or multi-modal machine learning, with direct experience in video understanding tasks.
Strong proficiency in Python and PyTorch, with solid software engineering fundamentals.
Hands-on experience with CLIP or similar vision-language/video embedding models for retrieval or representation learning.
Experience building or fine-tuning LLM-based systems for video/image understanding (e.g., captioning, video QA, summarization).
Familiarity with agentic system design — tool use, multi-step reasoning, and orchestration frameworks (e.g., LangChain, LlamaIndex, or custom agent loops).
Experience working with large-scale video data pipelines and vector search/retrieval infrastructure (e.g., FAISS, Milvus, or equivalent).
Preferred Qualifications
PhD with a research focus in video understanding, multi-modal learning, or vision-language models.
Experience with temporal action segmentation, action localization, or instruction-level video chunking algorithms.
Experience working with egocentric video datasets or head-mounted-device (HMD) captured data.
Track record of deploying production-scale video search or retrieval systems.
Experience integrating foundation or vision-language models (e.g., CLIP, VideoCLIP, RT-1/VLA variants) into perception or decision-making pipelines.
Publications in top-tier computer vision or ML venues (e.g., CVPR, ICCV, ECCV, NeurIPS, ICLR, etc).
Experience with humanoid robotics or embodied AI data pipelines is a plus.
Health insurance
Vision care
Dental coverage:
401(k)
Paid holidays
PTO (Paid Time Off)
Sick leave