... as image captioning and multimodal AI. Responsibilities : • Design and implement state-of-the-art vision-language models using deep learning frameworks. • Develop and fine-tune models that ...
... as image captioning and multimodal AI. Responsibilities : • Design and implement state-of-the-art vision-language models using deep learning frameworks. • Develop and fine-tune models that ...
Experience with image understanding tasks such as semantic segmentation, scene recognition, image captioning, visual question answering, image aesthetics, or image retrieval. Strong fundamental ...
Experience with image understanding tasks such as semantic segmentation, scene recognition, image captioning, visual question answering, image aesthetics, or image retrieval. Strong fundamental ...
Audio, image, or text applications - Source separation, text-to-speech, music synthesis, image segmentation, image captioning, question answering, language models, etc. Learn more about our ...
Audio, image, or text applications - Source separation, text-to-speech, music synthesis, image segmentation, image captioning, question answering, language models, etc. Learn more about our ...
... image captioning, question answering, language models, etc. Learn more about our innovative research: The Atlanta Area base salary range for this full-time position is $137,500-$168,200, which can ...
... image captioning, question answering, language models, etc. Learn more about our innovative research: The Atlanta Area base salary range for this full-time position is $137,500-$168,200, which can ...
Machine Learning Engineer -- Camera & Photos, Creative Foundations
San Diego, CA · On-site
$171.60 - $302.20/hr
Experience with image understanding tasks such as semantic segmentation, scene recognition, image captioning, visual question answering, image aesthetics, or image retrieval. * Strong fundamental ...
Machine Learning Engineer -- Camera & Photos, Creative Foundations
San Diego, CA · On-site
$171.60 - $302.20/hr
Experience with image understanding tasks such as semantic segmentation, scene recognition, image captioning, visual question answering, image aesthetics, or image retrieval. * Strong fundamental ...
Experience with image understanding tasks such as semantic segmentation, scene recognition, image captioning, visual question answering, image aesthetics, or image retrieval. Strong fundamental ...
Experience with image understanding tasks such as semantic segmentation, scene recognition, image captioning, visual question answering, image aesthetics, or image retrieval. Strong fundamental ...
Senior Audio AI Researcher
$88K - $120K/yr
Audio, image, or text applications - Source separation, text-to-speech, music synthesis, image segmentation, image captioning, question answering, language models, etc. Main Responsibilities
Senior Audio AI Researcher
$88K - $120K/yr
Audio, image, or text applications - Source separation, text-to-speech, music synthesis, image segmentation, image captioning, question answering, language models, etc. Main Responsibilities
Senior Audio AI Researcher
Atlanta, GA · On-site
$88K - $120K/yr
Audio, image, or text applications - Source separation, text-to-speech, music synthesis, image segmentation, image captioning, question answering, language models, etc. Main Responsibilities
Senior Audio AI Researcher
Atlanta, GA · On-site
$88K - $120K/yr
Audio, image, or text applications - Source separation, text-to-speech, music synthesis, image segmentation, image captioning, question answering, language models, etc. Main Responsibilities
Familiarity with captioning for large-scale data, including designing or applying automated captioning pipelines for image, video, and audio datasets. * Strong coding and prototyping ability in ...
Familiarity with captioning for large-scale data, including designing or applying automated captioning pipelines for image, video, and audio datasets. * Strong coding and prototyping ability in ...
Editing & Captioning: Assist the editing team with editing and captioning images from photographers on assignment in real time, as well as post assignment * Image Review & Management: Review images ...
Editing & Captioning: Assist the editing team with editing and captioning images from photographers on assignment in real time, as well as post assignment * Image Review & Management: Review images ...
Applied Scientist
San Jose, CA · On-site
Familiarity with captioning for large-scale data, including designing or applying automated captioning pipelines for image, video, and audio datasets. * Strong coding and prototyping ability in ...
Applied Scientist
San Jose, CA · On-site
Familiarity with captioning for large-scale data, including designing or applying automated captioning pipelines for image, video, and audio datasets. * Strong coding and prototyping ability in ...
Architect and optimize distributed multimodal inference pipelines for large-scale image, video, and audio captioning, tagging, and metadata generation. * Drive LLM/VLM inference optimization ...
Architect and optimize distributed multimodal inference pipelines for large-scale image, video, and audio captioning, tagging, and metadata generation. * Drive LLM/VLM inference optimization ...
Architect and optimize distributed multimodal inference pipelines for large-scale image, video, and audio captioning, tagging, and metadata generation. * Drive LLM/VLM inference optimization ...
Architect and optimize distributed multimodal inference pipelines for large-scale image, video, and audio captioning, tagging, and metadata generation. * Drive LLM/VLM inference optimization ...
Photographer (Part-Time)
New York, NY · On-site
Editing & Captioning: Assist the editing team with editing and captioning images from photographers on assignment in real time, as well as post assignment * Image Review & Management: Review images ...
Photographer (Part-Time)
New York, NY · On-site
Editing & Captioning: Assist the editing team with editing and captioning images from photographers on assignment in real time, as well as post assignment * Image Review & Management: Review images ...
Collaborate with Design Director to ensure moving-image work reflects the Walker's visual identity ... Familiarity with accessibility standards for digital video, including captioning and related ...
Collaborate with Design Director to ensure moving-image work reflects the Walker's visual identity ... Familiarity with accessibility standards for digital video, including captioning and related ...
Collaborate with Design Director to ensure moving-image work reflects the Walker's visual identity ... Familiarity with accessibility standards for digital video, including captioning and related ...
Collaborate with Design Director to ensure moving-image work reflects the Walker's visual identity ... Familiarity with accessibility standards for digital video, including captioning and related ...
Multimedia Producer
Minneapolis, MN · On-site
$62K/yr
Collaborate with Design Director to ensure moving-image work reflects the Walker's visual identity ... Familiarity with accessibility standards for digital video, including captioning and related ...
Multimedia Producer
Minneapolis, MN · On-site
$62K/yr
Collaborate with Design Director to ensure moving-image work reflects the Walker's visual identity ... Familiarity with accessibility standards for digital video, including captioning and related ...
... captioning, and in-depth data studies, particularly for visual and audio data. • Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality ...
... captioning, and in-depth data studies, particularly for visual and audio data. • Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality ...
... captioning, and in-depth data studies, particularly for visual and audio data. • Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality ...
... captioning, and in-depth data studies, particularly for visual and audio data. • Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality ...
Member of Technical Staff - Imagine Model
Palo Alto, CA · On-site
$180K - $440K/yr
... captioning, and in-depth data studies, particularly for visual and audio data. * Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality and ...
Member of Technical Staff - Imagine Model
Palo Alto, CA · On-site
$180K - $440K/yr
... captioning, and in-depth data studies, particularly for visual and audio data. * Design evaluation frameworks, metrics, benchmarks, evals, and reward models tailored to image/video/audio quality and ...
Image Captioning information
See salary details
$19.95 - $24.43
1% of jobs
$24.43 - $28.91
2% of jobs
$28.91 - $33.39
6% of jobs
$33.39 - $37.87
15% of jobs
$38.10 is the 25th percentile. Wages below this are outliers.
$37.87 - $42.35
21% of jobs
The median wage is $43.84 / hr.
$42.35 - $46.83
16% of jobs
$50.75 is the 75th percentile. Wages above this are outliers.
$46.83 - $51.31
17% of jobs
$51.31 - $55.79
10% of jobs
$55.79 - $60.27
4% of jobs
$60.27 - $64.75
6% of jobs
$64.75 - $69.23
2% of jobs
$19
$46
$69
How much do image captioning jobs pay per hour?
What is an image captioning job?
An Image Captioning job involves generating descriptive text for images using artificial intelligence or human expertise. Professionals in this field work with machine learning models, datasets, and natural language processing to create accurate and contextually relevant captions. This role is essential for improving accessibility, content organization, and searchability of visual media. It is commonly used in applications like social media, e-commerce, and automated reporting.
What are the typical responsibilities of someone working in image captioning?
Professionals in image captioning are primarily responsible for examining photos, graphics, or other visual data and crafting concise, accurate, and contextually appropriate captions. This process often involves using specialized software to annotate or tag images, ensuring consistency with style guidelines, and collaborating with editors, data teams, or project managers to align with project objectives. Daily tasks may also include reviewing and revising captions based on feedback, managing large batches of content, and maintaining organization within digital asset systems. The role is detail-oriented and can be performed individually or as part of a larger content or machine learning team depending on the employer.
What are the key skills and qualifications needed to thrive in the image captioning position, and why are they important?
To thrive in an Image Captioning role, you need strong attention to detail, language proficiency, and an ability to interpret visual content accurately. Familiarity with digital annotation tools, content management systems, or image labeling platforms is often required. Exceptional communication and time management skills help you handle large volumes of images and collaborate with team members or editors. These abilities ensure captions are clear, contextually relevant, and consistently meet quality and deadline standards.
Do image captioning jobs still exist?
What are the most commonly searched types of Image Captioning jobs?
The most popular types of Image Captioning jobs are:
What job categories do people searching Image Captioning jobs look for?
The top searched job categories for Image Captioning jobs are:

Vision Language Model Engineer
San Francisco, CA • On-site
Full-time
This job post has expired 1 day ago. Applications are no longer accepted.
Job description
EchoTwin AI is pioneering AI-driven infrastructure intelligence, redefining how cities are managed. As a Vision Language Model Engineer, you will design, develop, and optimize advanced vision-language models that integrate visual and textual data to enable intelligent systems, working closely with cross-functional teams to build applications such as image captioning and multimodal AI.
Responsibilities:
• Design and implement state-of-the-art vision-language models using deep learning frameworks.
• Develop and fine-tune models that combine computer vision and natural language processing for tasks like image captioning, visual question answering, and text-to-image generation.
• Collaborate with data scientists and software engineers to integrate models into production systems.
• Optimize model performance for accuracy, latency, and scalability in real-world applications.
• Conduct experiments to evaluate model performance and iterate on architectures and training pipelines.
• Stay up-to-date with the latest research in vision-language models and incorporate advancements into projects.
• Contribute to data preprocessing, augmentation, and annotation pipelines for multimodal datasets.
• Document model development processes and present findings to technical and non-technical stakeholders.
Qualifications:
Required:
• Bachelor’s, Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field (or equivalent experience).
• 3+ years of experience in machine learning, with a focus on vision-language models or multimodal AI.
• Hands-on experience with deep learning frameworks such as PyTorch or TensorFlow.
• Proven track record of building and deploying computer vision and/or NLP models.
• Proficiency in Python and relevant ML libraries (e.g., Hugging Face, OpenCV, Transformers).
• Experience with large-scale model training and optimization (e.g., distributed training, quantization).
• Strong understanding of neural network architectures (e.g., CNNs, Transformers, CLIP, or similar).
• Experience with multimodal datasets and preprocessing techniques for images and text.
• Familiarity with cloud platforms (e.g., AWS, GCP, Azure) and model deployment workflows.
• Strong problem-solving skills and ability to work in a fast-paced, collaborative environment.
• Excellent communication skills to explain complex technical concepts to diverse audiences.
Company:
The Physical AI Operating System for Real-Time Urban Intelligence | Physical AI at City Scale Founded in 2024, the company is headquartered in Boca Raton, USA, with a team of 11-50 employees. The company is currently Early Stage.