1

Video Annotation Jobs in New York (NOW HIRING)

... video creators, ranking systems, content strategy, content labeling and analysis * 1+ years working with GenAI products, prompt engineering, annotation, and/or content labeling and analysis

Driver Behavior Analyst

Manhattan, NY · On-site

$73K - $83K/yr

  • Medical

  • Dental

  • PTO

Behavioral Annotation: Analyze video and sensor data from US roads to label driver intent, pedestrian interactions, and complex maneuvers (e.g., unprotected left turns, merging in heavy traffic)

Showing results 21-35

Video Annotation information

See New York salary details

$28.4K

$65.4K

$103.9K

How much do video annotation jobs pay per year?

As of Aug 15, 2026, the average yearly pay for video annotation in New York is $65,410.00, according to ZipRecruiter salary data. Most workers in this role earn between $50,300.00 and $76,000.00 per year, depending on experience, location, and employer.

What is a video annotation?

A Video Annotation job involves labeling objects, activities, or events within videos to help train machine learning models. Annotators use specialized tools to draw bounding boxes, segment frames, or classify scenes to improve AI's ability to recognize visuals. This work is essential for applications like autonomous vehicles, facial recognition, and action recognition in AI systems.

What are the typical daily responsibilities of a video annotation specialist?

As a Video Annotation specialist, your daily tasks will generally involve watching video footage, identifying relevant objects or actions, and accurately labeling or tagging frames according to specific project guidelines. You may also review and validate annotations to ensure quality and consistency, collaborate with team members or project managers to clarify labeling instructions, and document any ambiguities or challenges encountered during annotation. Most roles are structured with clear targets or quotas for completed work, and you may work independently or as part of a larger team supporting AI development projects. The position requires strong concentration and the ability to handle repetitive tasks efficiently while maintaining high standards of accuracy.

What are the key skills and qualifications needed to thrive in the video annotation position, and why are they important?

To excel in Video Annotation, you need strong attention to detail, visual analysis skills, and familiarity with video processing concepts, often supported by a diploma or coursework in computer science or a related field. Knowledge of annotation tools such as CVAT, Labelbox, or VGG Image Annotator, and, in some cases, experience with basic scripting or data management platforms, is highly valued. Excellent focus, time management, and the ability to follow precise instructions help individuals stand out in this position. These abilities are crucial for ensuring the accuracy and quality of annotated video data, which directly impacts AI and machine learning model performance.

What are the most commonly searched types of Video Annotation jobs in New York?

The most popular types of Video Annotation jobs in New York are:

What are popular job titles related to Video Annotation jobs in New York?

For Video Annotation jobs in New York, the most frequently searched job titles are:

What job categories do people searching Video Annotation jobs in New York look for?

The top searched job categories for Video Annotation jobs in New York are:

Infographic showing various Video Annotation job openings in New York as of August 2026, with employment types broken down into 6% Internship, 49% Full Time, 20% Part Time, and 25% Contract. Highlights an 82% In-person, 6% Hybrid, and 12% Remote job distribution, with an average salary of $65,410 per year, or $31.4 per hour.

Computer Vision & Robotics Navigation Engineer

Kindredventures

Manhattan, NY • On-site

$90 - $130/hr

Other

Posted 9 days ago


Job description

About Mecka AI

Mecka AI is building the data infrastructure layer for robotics and embodied AI. We work with leading robotics companies and AI labs to collect, label, and validate the large-scale, real-world visual and spatial data used to train perception, manipulation, and control systems. Our work sits directly in the loop between raw sensor data, labeling pipelines, and deployed models.

About the Role

We are looking for a highly hands-on and product-oriented Computer Vision & Robotics Navigation Engineer to help us build the internal systems, algorithms, and tools that power our robotics data platform. Because this role involves hands-on testing, debugging, and local prototyping with our custom multi-sensor camera rigs, this is an on-site position based in our New York City office.

In this role, you will be the primary owner responsible for maintaining and improving our numerous SLAM and SfM systems across a variety of devices, including our custom camera rigs and iPhones. You will be deeply involved in the practical side of spatial computing—handling IMU noise modeling, sensor synchronization, and collaborating closely with our hardware team in China to ensure rigorous camera and sensor calibrations.

Beyond your core navigation focus, you will also act as the central hub for general computer vision support throughout the company. If you are passionate about multi-view geometry, enjoy building custom tooling to visualize complex trajectories, care deeply about data quality, and want your work to directly impact real robots—this role is for you.

What You’ll Do
  • Own the Navigation Systems: Act as the main engineer responsible for maintaining, optimizing, and improving our multiple SLAM and Structure from Motion (SfM) pipelines.

  • Sensor Calibration & Hardware Collaboration: Define, validate, and troubleshoot rigorous intrinsic and extrinsic calibration requirements for multi-camera setups and IMUs. You will communicate continuously with our hardware team in China—where the physical calibrations take place—while managing the algorithmic challenges of hardware-based SLAM locally, including temporal synchronization, rolling shutter correction, and IMU pre-integration.

  • Cross-Device Optimization: Ensure our spatial computing algorithms run robustly and accurately across a variety of hardware profiles, specifically our custom camera hardware and mobile devices (iOS/iPhone).

  • Company-Wide CV Support: Provide general computer vision expertise and support to various internal teams, assisting with pre- and post-processing, data validation, and automated labeling.

  • Design Internal Tooling: Ship custom tools (like Gradio or Rerun) to visualize images, video, 3D point clouds, and trajectories.

  • Debug & Inspect: Create interactive interfaces that help operations, annotators, and researchers inspect failure cases, understand edge conditions, and identify spatial labeling errors.

What We’re Looking For
  • Deep Navigation Expertise: A strong background in 3D computer vision and multi-view geometry, with proven experience building, maintaining, or improving SLAM, VIO, and SfM pipelines.

  • Practical SLAM & Calibration Skills: Deep knowledge of IMU kinematics (noise density, random walk biases) and rigorous camera calibration techniques (checkerboard/AprilTag targets, lens distortion models), with the ability to effectively communicate these technical requirements to cross-border hardware teams.

  • Hardware Familiarity: Experience working with spatial data from diverse hardware sources, such as custom camera rigs and mobile devices (iOS/iPhone).

  • Mathematical Fundamentals: An intuitive grasp of linear algebra, optimization, and the first principles of traditional CV and spatial tracking.

  • Engineering Rigor: A proven track record of software development expertise, consistently delivering high-quality, clean, efficient, and scalable code (especially in C++ and Python).

  • Adaptability: Comfortable iterating with users, bridging communication across time zones, supporting company-wide CV needs, and working alongside noisy, unstructured, real-world sensor data.

Strong Plus
  • Hands-on experience building CV/spatial tooling or apps such as dataset browsers, annotation tools, model debugging dashboards, or Gradio-style demos.

  • Experience with standard calibration and sensor fusion frameworks (e.g., Kalibr).

  • Exposure to ML infrastructure or data pipelines operating at scale.

Tech Stack
  • Python and C++ (Crucial for robust navigation/SLAM pipelines)

  • 3D Vision, Calibration & Optimization libraries (e.g., OpenCV, Ceres Solver, GTSAM, COLMAP, Kalibr)

  • PyTorch and deep learning CV libraries

  • Video and image processing pipelines (FFmpeg, etc.)

  • Internal web tooling and visualization (Rerun)

Note: The exact stack matters less than your ability to build, debug, and ship impactful spatial tools and algorithms.

What Success Looks Like
  • Our custom camera and iPhone SLAM/SfM systems perform reliably and efficiently under your ownership.

  • You establish a seamless feedback loop with the China hardware team, ensuring sensor rigs are tightly calibrated and trajectory estimates remain robust against real-world hardware noise.

  • Internal teams rely on your tools, navigation ground-truth, and general CV support daily.

  • Customers trust Mecka’s spatial data because the underlying algorithms and tooling are rock solid.

Who This Role Is Not For
  • Pure research roles with no production ownership.

  • Engineers looking for a remote or hybrid role—this requires physical presence with hardware testing in NYC.

  • Algorithm-only engineers who do not want to engage with the practical realities of hardware calibration, IMU noise, or cross-functional team communication.

Why This Role at Mecka?
  • Direct impact on how real robots are trained and navigate the world.

  • High ownership over core spatial data, multiple navigation systems, and CV support.

  • Close collaboration with leading robotics companies and AI labs.

  • The opportunity to build the multi-modal tooling layer that most teams wish they had.

#J-18808-Ljbffr