Humanoid robots

Custom Training Data for Humanoid Robots

Custom humanoid dataset capture on demand: bimanual, whole-body egocentric data recorded from a head-worn viewpoint, aligned across ten synchronized streams, and licensed to you.
Streams
Ten, synced under 2 ms
First batch
Median 48 h
Billing
Per validated hour
License
Buyer-owned
A contributor wearing a head-mounted camera beside a graphic of the first-person feed and annotation layers it produces.
Head-worn capture puts the camera where a humanoid actually sees from, hands entering frame from below.

Requirements

What data do humanoid robots need?

01

A head-worn viewpoint, not a third-person one

A humanoid perceives from a head-mounted camera, so custom humanoid dataset capture should be recorded from that same viewpoint: hands entering frame from below, occlusion during grasp, motion during locomotion.

02

Both hands, whole body, in frame

Bimanual coordination and whole-body balance are core to humanoid tasks. Capture needs two wrist cameras plus 3D hand and body pose so both the action and the supporting posture are recorded together.

03

Every stream on one clock

Ten synchronized streams per episode (4K head camera, two wrist cameras, depth, IMU, 3D hand and body pose, segmentation, and action labels), hardware-aligned under 2 ms.

04

Rights you can ship on

A perpetual, buyer-owned commercial license backed by a documented consent chain, so a policy trained on the data can go into a product.

Embodiment fit

Why does first-person human demonstration fit humanoid embodiment?

A humanoid with a head-mounted camera occupies almost exactly the viewpoint a head-worn rig records. Egocentric human capture matches that observation space far more closely than third-person video or single-arm teleoperation, which fix a different camera position and a different action space.

Whole-body pose alongside the hands means the data carries the balance and locomotion context a humanoid action depends on, not just the manipulation in isolation.

Bimanual capture

How does bimanual, whole-body capture work?

Two wrist cameras plus 3D hand pose keep both hands in frame with per-hand pose and visibility flags throughout coordinated, often asymmetric two-handed work. Body pose and IMU add the whole-body context: stance, reach, and transfer between locomotion and manipulation.

Instance segmentation and action segments mark what each hand is doing and when, so a policy can learn the handoff between the two.

The spec

How does a custom humanoid dataset spec work?

You describe the tasks your humanoid needs to perform. We turn it into a written spec covering tasks, environments, regions, capture tier, annotation depth, and accept criteria, then quote reach, timeline, and price against that document before any capture begins.

The collection runs through the same seven-stage end-to-end custom collection pipeline as every Firsthand program: contributor recruiting, capture, per-batch QA, anonymization, and licensed delivery. Rejected items are recaptured at our cost.

Fast · on demand

How fast does a first batch ship?

Where coverage is already deep, the first validated batch lands in a median of 48 hours from signed spec. Thin environments typically take three to four weeks, because vetting operators there is the bottleneck, and we quote that timeline up front.

On-demand top-ups run against the same spec and accept criteria after each training run, so new humanoid training data stays consistent with what you already hold.

Rights

What do licensing and consent look like?

Every custom collection ships under a perpetual, buyer-owned commercial license, and exclusivity is available so the data is never resold. Each episode links to a signed contributor release through a documented consent chain.

Faces, plates, and screens are detected and irreversibly blurred at ingest, and a reviewer confirms the result before anything ships. The anonymization record is delivered with the data.

Humanoid need vs delivery

What does a humanoid robotics team get, need by need?

Humanoid training data needs and what a Firsthand custom collection delivers for each.
Humanoid training needWhat Firsthand captures
Head camera viewpoint4K head-mounted camera matching a head-worn humanoid perception point
Two-hand manipulationTwo wrist cameras with per-hand visibility, capturing bimanual and asymmetric work
DepthSynchronized depth stream for scene and contact geometry
Whole-body motionIMU plus 3D body pose for stance, reach, and locomotion context
Hand pose3D hand pose per hand, aligned with the wrist camera streams
Action structureInstance segmentation, action segments, grasp and contact events
QualityMeasured against signed accept criteria; billed per validated hour; reject log with every batch
FormatsRLDS, LeRobot, WebDataset, plus HDF5, zarr, and Rerun .rrd

Every collection runs through the same seven-stage end-to-end custom collection pipeline.

FAQ

Custom training data for humanoid robots, answered.

What is custom training data for humanoid robots?
It is real-world, first-person human demonstration data captured against a written spec matched to a humanoid's embodiment: head camera viewpoint, two hands in frame, and whole-body motion. Firsthand handles contributors, capture, QA, consent, and delivery.
Why not use existing robot manipulation datasets instead?
Most existing datasets were scoped for a specific single-arm or third-person setup, not a humanoid embodiment. A custom collection is specified to your head camera and bimanual configuration and licensed to you, rather than shared across every buyer of an off-the-shelf set.
Does capture include whole-body pose, not just hands?
Yes. Each episode includes IMU and 3D body pose alongside 3D hand pose for both hands, so balance and locomotion context are recorded with the manipulation.
How is custom humanoid data billed?
Per validated hour, quoted against the written spec before any capture. You pay only for data that clears the accept criteria, and rejected items are recaptured at our cost.
Which delivery formats are supported?
RLDS, LeRobot, and WebDataset, plus HDF5, zarr, and Rerun .rrd. Converter source is readable, so you can retarget the schema to your own loader.
Can we start with a small pilot?
Yes. A Spec Pilot is a small, fast first batch that proves the spec and accept criteria before you scale volume on demand.

Sources

Where are the delivery formats documented?

Request a custom dataset.

Tell us the tasks your humanoid needs to perform. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.