Comparison

Human Demonstration Data vs Teleoperation Data

Egocentric human capture scales coverage cheaply; teleoperation grounds data in the exact robot embodiment.

Short answer

Teleoperation collects demonstrations by a human driving the robot, matching its embodiment exactly but limited by hardware, operator time, and site access. Human egocentric demonstration records people doing tasks directly with a worn camera rig — faster to scale and more diverse, at the cost of an embodiment gap that pairing or retargeting must close.

Side by side

Firsthand vs Teleoperation data, dimension by dimension.

DimensionFirsthandTeleoperation data
HoursCustom to your spec; 10,585 validated hours across the current catalog[VERIFY] — bound by robot and operator availability; typically far slower per hour than human capture
Capture setupHead + two wrist cameras, active-stereo depth, 3× IMU, hardware-triggered on one PTP clockHuman operator drives the robot (direct teleop, VR/motion-capture teleop, or a leader-follower rig)
StreamsTen synchronized streams incl. depth, IMU, 3D hand & body pose, segmentation, action segmentsRobot joint states and commands, plus whatever cameras the robot platform carries
SyncHardware-triggered; median 1.12 ms, 2.0 ms ceiling, drift written to metadata[VERIFY] — depends on the specific teleoperation rig and robot controller
Hand pose3D, handJoints[2][21][3], metric, in the head-camera frameNot applicable — the robot end-effector’s pose stands in for a hand
DepthMetric depth, 848×480 at 30 Hz, with per-pixel confidence[VERIFY] — depends on the robot’s onboard sensors, if any
Action labelsFrame-accurate action segments (verb, noun, ns timestamps)[VERIFY] — varies by lab; not standardized across teleoperation setups
Failure-case coverageTransit, occlusion and low-light shipped — not trimmed — with a published reject log[VERIFY] — bound by how many sites and operators the program can reach
LicensePerpetual, irrevocable, buyer-owned commercial license[VERIFY] — typically the collecting lab’s own data, license set by that lab
Commercial use allowedYes[VERIFY] — depends on the collecting organization’s license terms
Consent documentationWritten release per participant; consent artefacts + data card per batchNot applicable in the same sense — the operator, not a demonstration subject, is recorded
Delivery formatsRLDS, LeRobot, WebDataset, HDF5, zarr, Rerun .rrd[VERIFY] — varies by lab and robot stack
CostPer validated hour, quoted per engagement; free sample pack, no form[VERIFY] — includes robot hardware time, facility access, and operator labor per session

Third-party figures marked [VERIFY] are confirmed against the sources below before publishing. Firsthand figures are measured on-site.

When Teleoperation data is the right choice

  • The policy must control the exact robot embodiment you already have, with zero retargeting gap.
  • You need action data expressed directly in robot joint or end-effector commands, not human hand pose.
  • A small, tightly scoped skill set matters more than breadth of scenes, sites, or task variation.

When Firsthand is the right choice

  • You need broad coverage across sites, tasks, or geographies faster than a robot-and-operator fleet can reach on-demand.
  • The policy consumes hand pose, gaze, or first-person visual context that a robot’s own sensors don’t capture.
  • You want data to pair with existing teleoperation data for embodiment alignment rather than replace it outright.

FAQ

Common questions.

Is human demonstration data a replacement for teleoperation data?

Usually not a full replacement. The two are complementary: egocentric human data scales coverage cheaply and broadly, while teleoperation grounds a policy in the exact robot embodiment. Most programs blend both rather than choosing one.

What is the "embodiment gap" between human and robot demonstrations?

A human hand and a robot end-effector have different kinematics, degrees of freedom, and sensor placement. Data captured from a human body doesn’t map one-to-one onto robot joint commands, so a retargeting or alignment step is typically needed to close that gap.

Why does human demonstration data scale faster than teleoperation?

Teleoperation needs the physical robot, a controller rig, and an operator present for every session, which caps throughput to how many robot-hours a program can schedule. Human egocentric capture only needs a worn rig and a participant, so it scales across more sites and people at once.

What does Firsthand capture instead of teleoperation?

Firsthand captures human egocentric demonstration — ten synchronized streams including 3D hand pose, depth, and frame-accurate action segments, hardware-triggered on one clock — not robot teleoperation. It pairs with a program’s existing teleoperation data for embodiment alignment.

Does teleoperation data need consent documentation the way human capture does?

Teleoperation records an operator driving a robot, which is a different capture context than recording a demonstration subject performing a task. Firsthand’s consent chain — a written release per participant plus a data card per batch — applies to its own human egocentric capture.

Last reviewed: 2026-09-23

Judge it on your own skill spec.

Get the free 40-episode sample pack, or send us your target skills and hours for a scoped quote.