2

Entry Level Ai Media Captioning Jobs (NOW HIRING)

Not an entry-level role * Not a pure execution role long-term - this is a clear path to management ... We weave AI into everything we do , using the latest tech across all teams to innovate, work ...

Share AI Visibility Layer instructions after each article is published and confirm client ... Career Path: * Entry-level role with potential to grow into PR Manager or Client Success Manager.

... AI-powered tools for tasks such as image enhancement, asset generation, transcription, captioning ... Experience working in magazine, media, agency, or fast-paced content environments preferred.

WPP Media is WPP's AI-driven media operating unit, bringing together media, data, and partnerships ... Senior Associates are also typically responsible for managing entry-level team members. Key ...

... media, or agency environment (internship experience considered). * Strong proficiency in Adobe ... Familiarity with generative AI tools for editing, captioning, or effects. * Hands-on experience ...

Media Account Strategy Coordinator

Sandy Hook, CT · On-site

$20.50 - $26.75/hr

... Entry level candidates with previous internship experience are encouraged to apply. Ability to ... Exposure to and interest in using AI, particularly Claude, is a plus!Our office is located in Sandy ...

Showing results 41-60

Entry Level Ai Media Captioning information

See salary details

$11K

$111K

How much do entry level ai media captioning jobs pay per year?

As of Aug 20, 2026, the average yearly pay for entry level ai media captioning in the United States is $110,184.00, according to ZipRecruiter salary data. Most workers in this role earn between $110,000.00 and $110,500.00 per year, depending on experience, location, and employer.

What is an entry level AI media captioning job?

Entry level AI media captioning jobs involve creating, editing, or reviewing captions and subtitles for audio and video content using artificial intelligence tools. These roles typically require listening to media, correcting AI-generated captions, and ensuring text matches spoken dialogue accurately. No advanced experience is needed, but attention to detail, good language skills, and basic computer proficiency are important. Many positions are remote and may be part-time or project-based, providing a good starting point for those interested in media, technology, or accessibility services.

What are the key skills and qualifications needed to thrive as an entry level AI media captioning specialist?

To excel in Entry Level AI Media Captioning, you need strong language proficiency, attention to detail, and basic typing skills, often supported by a high school diploma or equivalent. Familiarity with captioning software, video editing tools, and speech recognition systems is typically required. Excellent listening abilities, time management, and the ability to work independently are valuable soft skills in this role. These competencies ensure captions are accurate, timely, and accessible, supporting content quality and inclusivity for diverse audiences.

What are some common challenges faced by entry level AI media captioning specialists, and how can I prepare for them?

Entry-level AI media captioning specialists often encounter challenges such as maintaining high accuracy while working with diverse audio qualities and accents, ensuring fast turnaround times, and adapting to evolving AI tools and software. To prepare, familiarize yourself with common captioning platforms, practice transcribing various audio types, and stay updated on industry standards for accessibility. Collaborating closely with editors and quality assurance teams can also help you improve your skills and adapt to feedback effectively.

What are the most commonly searched types of Ai Media Captioning jobs?

The most popular types of Ai Media Captioning jobs are:

Infographic showing various Entry Level Ai Media Captioning job openings in the United States as of June 2026, with employment types broken down into 58% Full Time, and 42% Part Time. Highlights an 70% Physical, 3% Hybrid, and 27% Remote job distribution, with an average salary of $110,184 per year, or $53 per hour.

Multimodal AI Researcher, Audio

Dolby Laboratories, Inc.

Atlanta, GA • On-site

Full-time

Re-posted 11 days ago


Job description

Join the leader in entertainment innovation and help us design the future. At Dolby, science meets art, and high tech means more than computer code. As a member of the Dolby team, you'll see and hear the results of your work everywhere, from movie theaters to smartphones. We continue to revolutionize how people create, deliver, and enjoy entertainment worldwide. To do that, we need the absolute best talent. We're big enough to give you all the resources you need, and small enough so you can make a real difference and earn recognition for your work. We offer a collegial culture, challenging projects, and excellent compensation and benefits, not to mention a Flex Work approach that is truly flexible to support where, when, and how you do your best work.

The Advanced Technology Group (ATG) is the research division of the company. ATG's mission is to look ahead, deliver insights, and innovate technological solutions that will fuel Dolby's continued growth. Our researchers have a broad range of expertise related to computer science and electrical engineering, such as AI/ML, algorithms, digital signal processing, audio engineering, image processing, computer vision, data science & analytics, distributed systems, cloud, edge & mobile computing, computer networking, and IoT.

Dolby is looking for a talented Senior Multimodal AI Researcher, Audio to join Dolby's research efforts and drive innovation in multimodal AI for audio applications, multimodal representations, and generative modeling for audio, speech, and music.  You will join the Machine Reasoning and Perception team to join a team of top-tier researchers working on challenging problems in multimodal AI for entertainment applications. You will focus on the creation and implementation of multimodal and audio AI technologies from the underlying theoretical concepts to the development of prototypes and demonstrations, with the goal to create new experiences.

You will drive key innovations for Dolby's core business which allow Dolby and its customers to build products that push the boundaries of sound and multimedia experiences.

Summary

You will push the boundaries of the state-of-the-art in audio and multimodal technologies. The ideal candidate would have a strong background in deep learning, both in terms of conceptual understanding, as well as practical experience, with previous exposure to audio applications. A core aspect of this role involves being able to keep up to date with the literature, implement, and innovate with the bleeding edge in generative models, self-supervised learning, and multi-modal learning.

With the explosion of large language models and natural language processing, you will partner closely with Dolby's worldwide AI research staff, which actively pursues the integration of such models into audio and media experiences. You will be able to hit the ground running, innovate, and contribute to such projects. Consequently, experience with language models, question answering, vision-language models, captioning, etc. would be highly beneficial.

We are looking for candidates with experience in any of the following:

  • Generative modeling for audio applications (diffusion models, autoregressive models, masked generative transformers).
  • Multimodal semantic understanding and multimodal reasoning.
  • Multimodal representations (audio-video, audio-text, audio-video-text).
  • Multimodal AI architectures, with a focus on generating audio, music, and speech (text-to-audio, video-to-audio, image-to-audio).
  • Self and semi-supervised learning.
  • AI driven audio enhancement, processing, and generation (for speech and music), such as speech enhancement and analysis, source separation, text-to-speech, text-to-music, music information retrieval, audio classification.
  • LLMs for audio applications.

What You Will Accomplish

  • Partner closely with other domain experts to refine and execute Dolby's technical strategy in artificial intelligence and machine learning.
  • Use deep learning to create new solutions (including foundation models) and enhance existing applications.
  • Push the state-of-the-art and develop intellectual property.
  • Transfer technology to product groups.
  • Establish research collaborations with external university partners.
  • Mentor interns on novel research problems.
  • Publish papers in top-tier conferences and journals.
  • Advise internal leaders on recent deep learning advancements in the industry and academia to further influence research direction and business decisions.

Key Requirements

  • Ph.D. in Computer Science or similar field.
  • A strong background in deep learning, both in terms of conceptual understanding, as well as practical experience.
  • Technical knowledge of audio fundamentals.
  • Deep passion for audio, music, and multimedia applications.
  • Deep knowledge on current machine learning literature.
  • Strong publication record, with publications in major machine learning conferences (e.g. NeurIPS, ICLR, ICML) or top domain-specific conferences is desirable (e.g., ACL, CVPR, ICASSP, Interspeech).
  • Highly skilled in Python and one or more popular deep learning frameworks (TensorFlow or PyTorch).
  • Ability to envision new technologies and turn them into innovative products.
  • Good communication and collaboration skills.

Learn more about our innovative research: https://www.dolby.com/about/innovation/empowering/

The Atlanta Area base salary range for this full-time position is $140,700-$170,000 , which can vary if outside this location, plus bonus, benefits, and some roles may also include equity. Our salary ranges are determined by role, level, and location. Within the range, individual pay is determined by work location and additional factors, including job-related skills, competencies, experience, market demands, internal parity, and relevant education or training. Your recruiter can share more about the specific salary range and perks and benefits for your location during the hiring process.

Dolby will consider qualified applicants with criminal histories in a manner consistent with the requirements of San Francisco Police Code, Article 49, and Administrative Code, Article 12

Equal Employment Opportunity:
Dolby is proud to be an equal opportunity employer. Our success depends on the combined skills and talents of all our employees. We are committed to making employment decisions without regard to race, religious creed, color, age, sex, sexual orientation, gender identity, national origin, religion, marital status, family status, medical condition, disability, military service, pregnancy, childbirth and related medical conditions or any other classification protected by federal, state, and local laws and ordinances.