Video annotation

Egocentric video annotation.

First-person video only trains a policy once it carries the right labels. We annotate the footage we capture — hand pose, action segments, contact events, and instance masks — against the same synchronized depth and hardware clock the video was recorded on, then deliver it in the format your training stack reads.
Labels
3D hand pose · action segments · contacts
Basis
Depth-aligned, not single-frame
Formats
RLDS · LeRobot · WebDataset
License
Buyer-owned
A contributor wearing a head-mounted capture rig performs a manual task, beside a graphic of synchronized hand-pose, depth, and action-segment annotation layers.
Annotation happens against the depth and clock the video was captured on, not guessed from a single frame.

Definition

What is egocentric video annotation?

Egocentric video annotation is labeling first-person (head- or wrist-mounted) video with the structured signals a robot policy trains on — 3D hand pose, action-segment boundaries, contact and grasp events, and object or instance masks — tied to the synchronized depth stream and hardware clock the video was recorded on, rather than a caption or single classification tag per clip.

How it works

What egocentric video annotation actually covers.

01

3D hand pose, per hand

Joint positions are annotated in 3D against the captured depth geometry, with a per-joint visibility flag, so occluded fingers are marked rather than guessed.

02

Action segments, not clip-level tags

Each task is broken into start- and end-timestamped action segments rather than one label for the whole clip, so a policy can learn where one action ends and the next begins.

03

Contact and grasp events

Hand-object contact onsets and grasp events are annotated against the depth stream, giving a policy the moment of physical interaction, not just proximity in the frame.

04

Instance segmentation for the objects in scope

Object and body masks are annotated where the program spec calls for them, so a policy can separate the hand, the task object, and the background.

At a glance

What you get.

Basis
Synchronized RGB + depth + IMU
Hand labels
3D joint pose · per-joint visibility
Action labels
Start/end-timestamped action segments
Interaction labels
Contact onsets · grasp events
Optional
Instance segmentation · de-identification
Delivery
RLDS · LeRobot · WebDataset · HDF5

Every collection runs through the same seven-stage end-to-end custom collection pipeline. End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline.

FAQ

Egocentric video annotation, answered.

What labels come with egocentric video annotation?
3D hand pose with per-joint visibility flags, start/end-timestamped action segments, and contact or grasp events as standard. Instance segmentation and de-identification (face and license-plate blurring) are added when the program spec calls for them.
How is hand pose annotated in 3D rather than guessed from the image?
Hand pose is annotated against the synchronized depth stream captured on the same hardware clock as the video, so joint positions reflect real geometry rather than being estimated from a single 2D frame.
Can annotation be scoped to a specific skill or vocabulary?
Yes. A skill spec sets the action vocabulary, environment, and label set before capture, so action-segment and contact labels use the terms your training stack expects rather than a generic taxonomy.
What format is annotated data delivered in?
RLDS, LeRobot, WebDataset, HDF5, or JSONL, matching the schema your training pipeline already reads, under a perpetual, buyer-owned license.
Is annotation available on video we already have, or only on new capture?
This program annotates video captured to your spec on our synchronized hardware, so hand pose and depth-based labels are grounded in the same clock the footage was recorded on.

Brief the gap. Get data back fast.

Tell us what your model is missing. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.