1

Image Captioning Jobs (NOW HIRING)

... image captioning, question answering, language models, etc. Learn more about our innovative research: The Atlanta Area base salary range for this full-time position is $137,500-$168,200, which can ...

Senior Audio AI Researcher

Atlanta, GA · On-site

$88K - $120K/yr

Audio, image, or text applications - Source separation, text-to-speech, music synthesis, image segmentation, image captioning, question answering, language models, etc. Main Responsibilities

Familiarity with captioning for large-scale data, including designing or applying automated captioning pipelines for image, video, and audio datasets. * Strong coding and prototyping ability in ...

Familiarity with captioning for large-scale data, including designing or applying automated captioning pipelines for image, video, and audio datasets. * Strong coding and prototyping ability in ...

Collaborate with Design Director to ensure moving-image work reflects the Walker's visual identity ... Familiarity with accessibility standards for digital video, including captioning and related ...

next page

Showing results 1-20

Image Captioning information

See salary details

$19

$46

$69

How much do image captioning jobs pay per hour?

As of Aug 27, 2026, the average hourly pay for image captioning in the United States is $46.80, according to ZipRecruiter salary data. Most workers in this role earn between $38.22 and $52.16 per hour, depending on experience, location, and employer.

What is an image captioning job?

An Image Captioning job involves generating descriptive text for images using artificial intelligence or human expertise. Professionals in this field work with machine learning models, datasets, and natural language processing to create accurate and contextually relevant captions. This role is essential for improving accessibility, content organization, and searchability of visual media. It is commonly used in applications like social media, e-commerce, and automated reporting.

What are the typical responsibilities of someone working in image captioning?

Professionals in image captioning are primarily responsible for examining photos, graphics, or other visual data and crafting concise, accurate, and contextually appropriate captions. This process often involves using specialized software to annotate or tag images, ensuring consistency with style guidelines, and collaborating with editors, data teams, or project managers to align with project objectives. Daily tasks may also include reviewing and revising captions based on feedback, managing large batches of content, and maintaining organization within digital asset systems. The role is detail-oriented and can be performed individually or as part of a larger content or machine learning team depending on the employer.

What are the key skills and qualifications needed to thrive in the image captioning position, and why are they important?

To thrive in an Image Captioning role, you need strong attention to detail, language proficiency, and an ability to interpret visual content accurately. Familiarity with digital annotation tools, content management systems, or image labeling platforms is often required. Exceptional communication and time management skills help you handle large volumes of images and collaborate with team members or editors. These abilities ensure captions are clear, contextually relevant, and consistently meet quality and deadline standards.

Do image captioning jobs still exist?

Yes, image captioning jobs still exist and are often part of roles in AI, machine learning, and data annotation. These jobs typically involve labeling images to improve computer vision systems and may require skills in image analysis and familiarity with annotation tools.
More about Image Captioning jobs

What are the most commonly searched types of Image Captioning jobs?

The most popular types of Image Captioning jobs are:

Infographic showing various Image Captioning job openings in the United States as of August 2026, with employment types broken down into 100% Full Time. Highlights an 100% In-person job distribution, with an average salary of $97,350 per year, or $46.8 per hour.

Vision Language Model Engineer

San Francisco, CA • On-site

Full-time

This job post has expired 1 day ago. Applications are no longer accepted.


Job description

Job Summary:
EchoTwin AI is pioneering AI-driven infrastructure intelligence, redefining how cities are managed. As a Vision Language Model Engineer, you will design, develop, and optimize advanced vision-language models that integrate visual and textual data to enable intelligent systems, working closely with cross-functional teams to build applications such as image captioning and multimodal AI.
Responsibilities:
• Design and implement state-of-the-art vision-language models using deep learning frameworks.
• Develop and fine-tune models that combine computer vision and natural language processing for tasks like image captioning, visual question answering, and text-to-image generation.
• Collaborate with data scientists and software engineers to integrate models into production systems.
• Optimize model performance for accuracy, latency, and scalability in real-world applications.
• Conduct experiments to evaluate model performance and iterate on architectures and training pipelines.
• Stay up-to-date with the latest research in vision-language models and incorporate advancements into projects.
• Contribute to data preprocessing, augmentation, and annotation pipelines for multimodal datasets.
• Document model development processes and present findings to technical and non-technical stakeholders.
Qualifications:
Required:
• Bachelor’s, Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field (or equivalent experience).
• 3+ years of experience in machine learning, with a focus on vision-language models or multimodal AI.
• Hands-on experience with deep learning frameworks such as PyTorch or TensorFlow.
• Proven track record of building and deploying computer vision and/or NLP models.
• Proficiency in Python and relevant ML libraries (e.g., Hugging Face, OpenCV, Transformers).
• Experience with large-scale model training and optimization (e.g., distributed training, quantization).
• Strong understanding of neural network architectures (e.g., CNNs, Transformers, CLIP, or similar).
• Experience with multimodal datasets and preprocessing techniques for images and text.
• Familiarity with cloud platforms (e.g., AWS, GCP, Azure) and model deployment workflows.
• Strong problem-solving skills and ability to work in a fast-paced, collaborative environment.
• Excellent communication skills to explain complex technical concepts to diverse audiences.
Company:
The Physical AI Operating System for Real-Time Urban Intelligence | Physical AI at City Scale Founded in 2024, the company is headquartered in Boca Raton, USA, with a team of 11-50 employees. The company is currently Early Stage.