Egocentric video

Egocentric video data collection, first-person and to spec.

First-person video of real people doing real tasks, from a phone in a chest mount up to a full synchronized multi-camera rig. Same spec, same QA bar, whichever capture tier the project needs.
Capture tiers
Phone · head rig · multi-cam
Contributors
Paid hourly
First batch
Median 48 h
Anonymized
Faces · plates · screens
A contributor films a kitchen task hands-free with a phone in a chest mount, beside the annotated first-person frame it produces.
Phone-tier egocentric capture: hands-free, to spec, with the same annotation layers added downstream.

Definition

What is egocentric video data collection?

Egocentric video data collection is recording video from the point of view of the person doing a task, using a head, chest, or glasses-mounted camera. It captures hands, tools, and objects the way the actor sees them, which is the viewpoint embodied AI and assistive models need to learn from.

How it works

Three capture tiers, one standard.

01

Phone tier

A contributor films hands-free on a phone in a chest mount. It is the fastest way to reach broad geographic and household diversity.

02

Head-rig tier

A head-mounted camera matches the eye-line viewpoint and adds IMU, for tasks where gaze and head motion matter.

03

Multi-camera tier

Synchronized head, wrist, and depth streams on one hardware clock, for manipulation work that needs 3D hand pose and contacts.

04

Anonymized before delivery

Faces, plates, and screens are detected and irreversibly blurred at ingest and confirmed by a reviewer before anything ships.

At a glance

What you get.

Tiers
Phone · head rig · multi-camera
Annotation
Actions · hand pose · objects · transcripts
Anonymization
Faces · plates · screens, human-reviewed
Spec to first batch
a median of 48 hours (deep coverage)
Formats
RLDS · LeRobot · WebDataset · MP4 + JSONL

Every collection runs through the same seven-stage end-to-end custom collection pipeline. End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline.

FAQ

Egocentric video data collection, answered.

Which capture tier should I choose?
Phone tier for breadth and diversity, head-rig tier when gaze matters, and multi-camera tier when your model needs synchronized 3D hand pose. Many programs mix tiers under one spec.
How is contributor privacy protected?
Every contributor signs a documented release, bystanders and screens are blurred at ingest, and the anonymization record ships with the data.
Can I see sample egocentric footage first?
Yes. Raw clips are published on the examples page, and a downloadable sample pack is available from the homepage.

Brief the gap. Get data back fast.

Tell us what your model is missing. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.