Guide

Egocentric data for humanoid robots: what to collect

First-person, bimanual demonstration with 3D hand pose maps cleanly to humanoids.
By the Firsthand capture teamLast updated September 23, 2026

Short answer

Humanoids perceive from a head camera and act with two hands, so their training data should too. Egocentric human capture — head plus dual wrist cameras, depth, and 3D bimanual hand pose — matches a humanoid observation and action space more closely than any third-person or single-arm source.

Why does egocentric human data fit humanoids?

A humanoid with a head-mounted camera occupies almost exactly the viewpoint a head-worn rig records: hands entering frame from below, occlusion during grasp, motion during locomotion. Bimanual 3D hand pose gives the two-arm action structure a humanoid needs, which single-arm teleoperation data cannot supply.

What should you collect?

  • Bimanual tasks with coordinated, asymmetric roles (one hand stabilises, one acts)
  • 3D hand pose for both hands, with a per-joint visibility flag
  • Whole-body motion where locomotion and manipulation overlap
  • Failure and recovery — dropped objects, regrasps, mid-task correction

Egocentric human data pairs with your teleoperation data: it scales coverage cheaply, teleop grounds it in the exact robot.

Check it against the sample pack.

40 episodes across 4 environments, delivered in the exact schema these guides describe.