3D hand pose, per hand
Joint positions are annotated in 3D against the captured depth geometry, with a per-joint visibility flag, so occluded fingers are marked rather than guessed.
Video annotation

Definition
Egocentric video annotation is labeling first-person (head- or wrist-mounted) video with the structured signals a robot policy trains on — 3D hand pose, action-segment boundaries, contact and grasp events, and object or instance masks — tied to the synchronized depth stream and hardware clock the video was recorded on, rather than a caption or single classification tag per clip.
How it works
Joint positions are annotated in 3D against the captured depth geometry, with a per-joint visibility flag, so occluded fingers are marked rather than guessed.
Each task is broken into start- and end-timestamped action segments rather than one label for the whole clip, so a policy can learn where one action ends and the next begins.
Hand-object contact onsets and grasp events are annotated against the depth stream, giving a policy the moment of physical interaction, not just proximity in the frame.
Object and body masks are annotated where the program spec calls for them, so a policy can separate the hand, the task object, and the background.
At a glance
Every collection runs through the same seven-stage end-to-end custom collection pipeline. End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline.
FAQ
Tell us what your model is missing. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.