3D hand pose, not a 2D box around the hand
Per-frame hand keypoints are tracked in 3D against depth, with a visibility flag per joint, so occluded fingers during a grasp are marked rather than guessed.
Manipulation signal

Definition
Hand-object interaction video data is footage of a hand approaching, contacting, and manipulating an object, captured with 3D hand pose and contact-event annotation synchronized to a depth stream on one hardware clock, so a manipulation policy can learn the approach, grasp, and release rather than just seeing a hand near an object in a single frame.
How it works
Per-frame hand keypoints are tracked in 3D against depth, with a visibility flag per joint, so occluded fingers during a grasp are marked rather than guessed.
When and where fingers make and break contact is annotated against the same clock as the video and depth — the signal a manipulation policy actually conditions on.
Metric depth aligned to the hand and object gives a policy the approach vector and distance, not just a 2D silhouette of the interaction.
A wrist-mounted camera adds a close-range view of the interaction alongside the head view, useful when the grasp detail matters more than the full-arm reach.
At a glance
Every collection runs through the same seven-stage end-to-end custom collection pipeline. End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline.
FAQ
Tell us what your model is missing. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.