Human activity

Human activity video data collection.

Robots and embodied AI models learn manipulation and everyday tasks by watching how people actually do them. We capture that human activity from a first-person view, with hand pose and action labels synchronized to the video, rather than a third-person security-camera angle.
Viewpoint
First-person, egocentric
Capture tiers
Phone · head rig · multi-cam
Labels
Hand pose · action segments · contact
License
Buyer-owned
A contributor films a kitchen task hands-free with a phone in a chest mount, beside the annotated first-person frame it produces.
Human activity captured from the actor’s own point of view, with hand pose and action labels added downstream.

Definition

What is human activity video data?

Human activity video data is footage of people performing real tasks, used to train models that need to recognize or imitate what a body does. For robot and embodied AI training this works best captured egocentrically, from a head, chest, or wrist-mounted camera, so the viewpoint and hand motion match what the model will see and act on.

How it works

Why egocentric capture fits robot training better than third-person footage.

01

The viewpoint has to match the task

A robot arm or humanoid perceives from its own camera, not an overhead security angle. Egocentric capture puts the camera where the actor’s eyes or wrist would be, so the geometry matches what the policy will see.

02

Hands and objects need to stay in frame

Third-person shots lose the hand-object relationship whenever the actor turns or reaches. A head-rig or wrist camera keeps the manipulation in view for the whole action.

03

Activity needs to be labeled, not just recorded

Training on activity means knowing what happened and when — action segments, hand pose, and contact events synchronized to the video, not a single tag per clip.

04

Coverage has to span real variation

People do the same task differently. A spec sets the skill and environment, then contributors perform it across the range a policy actually needs to generalize over.

At a glance

What you get.

Capture
Head-mounted, wrist, or phone first-person
Sensors
RGB · depth · IMU
Labels
Hand pose · action segments · contact
Spec to first batch
a median of 48 hours (deep coverage)
Delivery
RLDS · LeRobot · WebDataset · HDF5

Every collection runs through the same seven-stage end-to-end custom collection pipeline. End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline.

FAQ

Human activity video data, answered.

Is human activity video data the same as action recognition data?
They overlap. Action recognition typically needs only a clip-level label, while robot and embodied AI training also needs synchronized hand pose and action-segment boundaries so a policy can learn the motion, not just classify the clip.
Why capture human activity egocentrically instead of third-person?
A robot or humanoid acts from its own viewpoint, not a fixed camera across the room. First-person capture matches that geometry directly, and keeps the hands and manipulated object in frame through the whole action.
What annotation ships with human activity video?
Action segments, 3D hand pose with per-joint visibility, and contact events, synchronized to the video on one hardware clock, plus anonymization of faces, plates, and screens before delivery.
Can activity data be scoped to specific tasks or environments?
Yes. A spec sets the skill, environment, and capture tier, then contributors perform that task. Bimanual manipulation programs are a common starting point for hands-on activity.
How is human activity video data priced and delivered?
Per validated hour, quoted against your spec before capture. Delivery is in the format your training stack already reads — RLDS, LeRobot, WebDataset, or HDF5 — under a perpetual, buyer-owned license.

Brief the gap. Get data back fast.

Tell us what your model is missing. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.