1

Video Annotation Jobs in Atlanta, GA (NOW HIRING)

Monitor live video feeds and robot telemetry * Perform real-time movement adjustments and task ... Support data annotation and quality validation activities * Maintain accurate operational records ...

Monitor live video feeds and robot telemetry * Perform real-time movement adjustments and task ... Support data annotation and quality validation activities * Maintain accurate operational records ...

Video Annotation information

See Atlanta, GA salary details

$25K

$57.5K

$91.4K

How much do video annotation jobs pay per year?

As of Aug 24, 2026, the average yearly pay for video annotation in Atlanta, GA is $57,495.00, according to ZipRecruiter salary data. Most workers in this role earn between $44,200.00 and $66,800.00 per year, depending on experience, location, and employer.

What is a video annotation?

A Video Annotation job involves labeling objects, activities, or events within videos to help train machine learning models. Annotators use specialized tools to draw bounding boxes, segment frames, or classify scenes to improve AI's ability to recognize visuals. This work is essential for applications like autonomous vehicles, facial recognition, and action recognition in AI systems.

What are the typical daily responsibilities of a video annotation specialist?

As a Video Annotation specialist, your daily tasks will generally involve watching video footage, identifying relevant objects or actions, and accurately labeling or tagging frames according to specific project guidelines. You may also review and validate annotations to ensure quality and consistency, collaborate with team members or project managers to clarify labeling instructions, and document any ambiguities or challenges encountered during annotation. Most roles are structured with clear targets or quotas for completed work, and you may work independently or as part of a larger team supporting AI development projects. The position requires strong concentration and the ability to handle repetitive tasks efficiently while maintaining high standards of accuracy.

What are the key skills and qualifications needed to thrive in the video annotation position, and why are they important?

To excel in Video Annotation, you need strong attention to detail, visual analysis skills, and familiarity with video processing concepts, often supported by a diploma or coursework in computer science or a related field. Knowledge of annotation tools such as CVAT, Labelbox, or VGG Image Annotator, and, in some cases, experience with basic scripting or data management platforms, is highly valued. Excellent focus, time management, and the ability to follow precise instructions help individuals stand out in this position. These abilities are crucial for ensuring the accuracy and quality of annotated video data, which directly impacts AI and machine learning model performance.

What are the most commonly searched types of Video Annotation jobs in Atlanta, GA?

The most popular types of Video Annotation jobs in Atlanta, GA are:

What are popular job titles related to Video Annotation jobs in Atlanta, GA?

For Video Annotation jobs in Atlanta, GA, the most frequently searched job titles are:

What job categories do people searching Video Annotation jobs in Atlanta, GA look for?

The top searched job categories for Video Annotation jobs in Atlanta, GA are:

Infographic showing various Video Annotation job openings in Atlanta, GA as of August 2026, with employment types broken down into 7% Internship, 56% Full Time, 14% Part Time, and 23% Contract. Highlights an 75% In-person, 10% Hybrid, and 15% Remote job distribution, with an average salary of $57,495 per year, or $27.6 per hour.

Senior Computer Vision Engineer (Egocentric), Data Foundry Software

Front Door Defense

Atlanta, GA • On-site

$160 - $230/hr

Other

Posted 10 days ago


Job description

Senior Computer Vision Engineer (Egocentric), Data Foundry Build and scale an egocentric perception stack from data collection to production deployment

Location: Atlanta, Georgia

About The Role Computer Vision Engineer And Technologist

Stord operates the largest independent e-commerce fulfillment network in the US — 20+ fulfillment centers, 4,000+ warehouse associates, and nearly 100 million packages shipped annually. We are building a new business line that turns this operational infrastructure into some of the most valuable training data assets in physical AI. We are looking for an experienced computer vision engineer and technologist to build and scale this business from the ground up.

What You Will Own

You will own the early egocentric video stack — data collection, vision models and pipelines, and rigs. You'll partner closely with a small team to operationalize. This is a builder-operator role. You will:

  • Define and deliver the product. You will own the data product across quality tiers — from RGB egocentric video through depth-enhanced and full multimodal capture with hand pose and annotations. You will decide what gets built, in what order, based on what buyers will actually pay for. You will hold the line on quality.
  • Run the capture and delivery program. You will stand up the warehouse capture operation: camera and rig hardware selection, enrollment, edge processing, and the processing pipelines that package datasets for delivery. You will coordinate across warehouse operations, engineering, and customers to ship datasets on spec and on schedule.
  • Build the perception stack. Detection, tracking, and segmentation, plus depth/3D reconstruction and 6DoF, multi-view 3D hand/body pose estimation from egocentric and fixed-camera capture.
  • Stand up VLM-assisted and automated labeling with human-in-the-loop QA to drive down cost per annotated hour, and integrate the annotation tooling.
  • Own the hardware vision intersection. Camera calibration, epipolar/multi-view geometry, and frame-accurate time-sync across multi-camera and egocentric rigs; derive 3D pose by triangulation where no direct sensor exists.
  • Train and ship models. Design, fine-tune, and optimize CV/multimodal models on large unstructured video datasets, and get them reproducible and production-ready, not stuck in a notebook.
What You Bring
  • Experiencing standing up and scaling an egocentric perception stack. You have built and run a similar product end to end at a robotics or AI data company. You have driven the full lifecycle: hardware setup, embedded perception, data pipelines, ensuring quality, and delivering it to production teams who depend on it.
  • 8+ years building and shipping production computer-vision/perception systems (or an MS/PhD in CV, ML, or robotics plus 6+ years hands-on), including systems that ran on messy real-world data, not just benchmarks.
  • Deep expertise in computer vision and tooling — track record of leveraging existing tooling and designing, training, and debugging CNNs and vision transformers from scratch.
  • Strong command of geometric computer vision: camera calibration, depth estimation, and 2D/3D pose estimation
  • End-to-end ownership of a major perception problem: from data and model design through evaluation, optimization, and deployment, with measurable accuracy and reliability outcomes.
  • Track record of setting technical direction for a team or large workstream and raising the bar for other engineers.
  • Proven ability to take ambiguous, 0→1 problems with no established playbook and drive them to a working system with limited resources.
  • Experience with large unstructured datasets (video/multimodal) and the eval discipline to instrument accuracy rather than eyeball it.
  • Expert Python and strong software-engineering fundamentals; C++ where performance demands it.
Why This Role

This is a rare opportunity to build a high-growth business from the ground-up with infrastructure and resources to support. You will have:

  • A structural moat that no startup can replicate
  • Direct access to the fastest-growing buyer market in AI
  • CTO/Co-Founder as your direct partner.
#J-18808-Ljbffr