Manipulation signal

Hand-object interaction video data.

A manipulation policy learns from the moment fingers meet an object, not from a caption describing the task. We capture hand-object interaction on a synchronized rig — 3D hand pose, contact events, and depth aligned to one hardware clock — so the exact instant of contact is in the data, not guessed after the fact.
Signals
Hand pose · contact events · depth
Basis
Depth-aligned, per-frame
Formats
RLDS · LeRobot · WebDataset
License
Buyer-owned
A human hand grasps a household object mid-demonstration, with a depth sensor and annotated contact point visible.
Contact is the signal a grasp is made of. It only exists in the data if it is captured at the moment it happens.

Definition

What is hand-object interaction video data?

Hand-object interaction video data is footage of a hand approaching, contacting, and manipulating an object, captured with 3D hand pose and contact-event annotation synchronized to a depth stream on one hardware clock, so a manipulation policy can learn the approach, grasp, and release rather than just seeing a hand near an object in a single frame.

How it works

What makes interaction data usable for a manipulation policy.

01

3D hand pose, not a 2D box around the hand

Per-frame hand keypoints are tracked in 3D against depth, with a visibility flag per joint, so occluded fingers during a grasp are marked rather than guessed.

02

Contact events at the moment they happen

When and where fingers make and break contact is annotated against the same clock as the video and depth — the signal a manipulation policy actually conditions on.

03

Depth for approach and collision reasoning

Metric depth aligned to the hand and object gives a policy the approach vector and distance, not just a 2D silhouette of the interaction.

04

Optional wrist view for fine manipulation

A wrist-mounted camera adds a close-range view of the interaction alongside the head view, useful when the grasp detail matters more than the full-arm reach.

At a glance

What you get.

Basis
Synchronized RGB + depth + IMU
Hand labels
3D joint pose · per-joint visibility
Interaction labels
Contact onsets · grasp events
Sync
Hardware-triggered, one clock domain
Optional
Wrist camera · instance segmentation
Delivery
RLDS · LeRobot · WebDataset · HDF5

Every collection runs through the same seven-stage end-to-end custom collection pipeline. End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline.

FAQ

Hand-object interaction data, answered.

How is hand-object contact captured, rather than estimated after the fact?
Contact onsets and grasp events are annotated against the same synchronized depth stream and hardware clock the video was recorded on, so the moment of contact reflects real geometry rather than being inferred from a single 2D frame.
What is the difference between hand pose data and hand-object interaction data?
Hand pose data tracks the hand’s joints. Hand-object interaction data adds the contact and grasp events between the hand and the object it is manipulating, synchronized to the same depth and clock, which is what a manipulation policy conditions on.
Can this be captured with a wrist camera as well as a head-mounted view?
Yes. A wrist-mounted camera can run alongside the head view for a close-range view of the grasp, useful for fine manipulation where the head rig alone would lose the detail.
What formats is hand-object interaction data delivered in?
RLDS and LeRobot for policy training, plus WebDataset and HDF5, with per-stream nanosecond timestamps and calibration, under a perpetual, buyer-owned license.
Can interaction data be scoped to specific objects or grasps?
Yes. A spec names the objects, grasps, and environments the policy needs to generalize across, then capture and annotation are built around that spec rather than a generic taxonomy.

Brief the gap. Get data back fast.

Tell us what your model is missing. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.