Guide

How to Collect Egocentric Video Data

A methodology walkthrough of how egocentric video data collection actually works, from spec to delivery — for teams evaluating how to source or commission it.
By the Firsthand capture teamLast updated September 1, 2026

Short answer

Collecting egocentric video data means defining a capture spec (streams, sync ceiling, coverage), recruiting consented contributors to film first-person footage against that spec, hardware-synchronizing the sensor streams to a shared clock, validating every hour against the sync ceiling before it counts, anonymizing faces and identifying details, and delivering it in a training-ready format.

Last reviewed:

Why does the collection method matter, not just the footage?

Two datasets can look similar as video and still differ enormously in training value, depending on whether the streams are actually synchronized, whether every hour was validated before being counted, and whether consent and provenance are documented. The method behind the collection — not just the final clips — is what determines whether a dataset is usable for training a policy.

This page describes the methodology. For a productized service that runs this process for a project, see egocentric video data collection.

How is a collection program planned before any filming starts?

Collection starts from a capture spec: which sensor streams are needed (video, depth, pose), what sync-error ceiling the data must meet, what tasks and environments need coverage, and whether failure cases and near-misses are required alongside successful attempts. Defining this up front is what keeps the resulting dataset checkable against a standard rather than assessed after the fact.

How is the footage actually captured?

Paid, consented contributors film first-person footage of real tasks using a capture tier that matches the spec — a phone in a chest mount for broad, fast diversity, a head-mounted rig, or a full multi-camera setup with wrist cameras and depth for manipulation detail.

The capture tier is chosen to match the spec, not the other way around.
Capture tierWhat it capturesBest for
Phone (chest mount)Hands-free first-person video onlyFast, broad household and geographic diversity
Head-mounted rigFirst-person video, closer to eye-level viewpointTasks where head orientation and gaze matter
Multi-camera rigHead-mounted video plus wrist cameras and depthManipulation detail, occlusion during contact

The capture tier is chosen to match the spec, not the other way around.

How are the sensor streams kept synchronized during capture?

Video, depth, and pose are hardware-triggered off a shared clock using PTP (IEEE 1588) rather than software timestamps, which drift. This is what keeps cross-stream sync error low and measurable instead of assumed from the hardware spec.

How is footage validated before it counts as delivered data?

Every hour of footage is checked against the sync-error ceiling and the capture spec before it is counted as a validated hour. Footage that fails the check is rejected and, where the spec calls for it, recaptured — a raw hour of recording and a validated hour are not the same thing.

How is footage anonymized before delivery?

Faces, license plates, and screens are detected and blurred as part of the pipeline, and every participant’s consent is recorded and tied to the specific footage they appear in ����� so a buyer can trace exactly what was agreed to for any clip in the delivered dataset.

What format is the data delivered in?

Delivery uses standard robot-learning formats such as RLDS or LeRobot, with a buyer-owned, perpetual license and a documented consent chain shipping with every batch rather than being assembled after the fact.

FAQ

How to Collect Egocentric Video Data, answered.

01

What is the first step in collecting egocentric video data?

Defining a capture spec: which sensor streams are needed, what sync-error ceiling the data must meet, what tasks and environments require coverage, and whether failure cases need to be included alongside successful attempts.

02

How is camera synchronization achieved during collection?

Video, depth, and pose streams are hardware-triggered off a shared clock using PTP (IEEE 1588), rather than relying on software timestamps that drift over a session.

03

What makes an hour of footage "validated" rather than just recorded?

A validated hour has passed a check against the sync-error ceiling and the capture spec before it counts toward delivery. Footage that fails that check is rejected, and recaptured where the spec requires it.

04

How is privacy handled during egocentric data collection?

Faces, license plates, and screens are detected and blurred, and every participant’s consent is recorded and tied to the specific footage they appear in, so a buyer can trace what was agreed to for any clip.

05

What is the difference between this methodology and a commissioned collection service?

This page describes how the collection process works end to end. A commissioned service runs that process against a project’s specific spec and delivers the resulting dataset — see egocentric video data collection for that service.

Sources

Where is this documented?

Check it against the sample pack.

40 episodes across 4 environments, delivered in the exact schema these guides describe.