Egocentric capture · embodied AI
First-person video that survives contact with a policy.
Egocentric capture for embodied AI. Every clip ships with depth, hand pose, and body pose — synchronized to under two milliseconds, with a documented consent chain and a commercial license attached.
- Sync ceiling
- 2.0 ms, hardware-triggered
- License
- Commercial, buyer-owned
- Enrichment
- Depth · hands · body · masks
- Delivery
- RLDS · LeRobot · WebDataset
Synthetic frame, real schema. Toggle any layer — every one of these ships on every clip.
Validated hours delivered
how we count ›
Hours that passed sync + QA and shipped to a buyer. Recaptured and rejected hours excluded.
Median cross-stream sync error
how we count ›
Median absolute timestamp offset between head camera and every other stream, over 12,480 measured frames.
Documented consent
how we count ›
Every participant signs a written commercial release before capture. Bystanders are de-identified or the clip is discarded.
Median time to first delivery
how we count ›
Signed spec to first validated batch in your hands. Measured across Spec Pilot engagements.
The sample
See the data before you talk to anyone.
A 30-second slice with nothing stripped out: same schema, same enrichment layers, same license terms as a production batch. No form, no call, no NDA.
Rendered from the schema below. Toggle layers, scrub the playhead, then read the files it came from.
Head camera, 30 s, 3840×2160 · 60 fps
download ↓Hand + body joints, masks, action segments, calibration
download ↓Rerun recording — every stream on one timeline
download ↓Download the full 2.4 GB pack
40 episodes across four environments, all streams at native rate. Takes an email — asked for after the free files above, not before.
What ships with every clip
Ten streams, one clock, one episode file.
Not an à la carte menu. Every validated hour we deliver carries the whole stack, in the schema below, whether you asked for depth or not.
| Stream | Spec | Detail |
|---|---|---|
| Head camera | 3840 × 2160 · 60 fps | Global-shutter, 120° HFOV, rolling-shutter-free, H.265 + lossless keyframes |
| Wrist cameras × 2 | 1920 × 1080 · 60 fps | Left + right, 100° HFOV, rigid mount, extrinsics re-solved per session |
| Depth | 848 × 480 · 30 fps | Active stereo, 0.3–4 m range, metric millimetres, per-frame confidence map |
| IMU | 200 Hz · 6-DoF | Accel + gyro, bias-calibrated, hardware-timestamped on the same clock domain |
| Hand pose | 21 keypoints × 2 hands | 3D metric, per-joint visibility flag, contact events on grasp and release |
| Body pose | 24 joints | 3D, root-relative and world-frame, torso and forearm chains resolved |
| Segmentation | Instance masks | Manipulated objects + target surfaces, tracked IDs across the episode |
| Action segments | Verb + noun taxonomy | 97 verbs, 512 nouns, start/end to the frame, human-reviewed |
| Formats | RLDS · LeRobot · WebDataset | Also HDF5, zarr, and .rrd for Rerun. Converters shipped as source. |
| License | Commercial · buyer-owned | Perpetual, irrevocable, model-weights-clean. Exclusivity available. |
Colour dots map to the modality legend used in every chart on this site. Full field-level schema in the episode schema docs.
How we capture
One clock domain, or it isn't a dataset.
Ten streams recorded on separate sensors are ten separate videos until something forces them onto the same timebase. We hardware-trigger every sensor from a single PTP grandmaster, and we publish the resulting error distribution.


| Head camera | Sony IMX585 global shutter · 120° M12 | Carbon headband, 42 g |
| Wrist cameras | IMX290 ×2 · 100° M12 | Velcro cuff, rigid plate |
| Depth | Intel RealSense D435if | Forehead, 18 mm below head cam |
| IMU | Bosch BMI088 ×3 · 200 Hz | Head, both wrists |
| Clock | PTP grandmaster + hardware trigger | Belt pack, 190 g |
| Storage | 2 TB NVMe · 6 h continuous | Belt pack |
Rig architecture
Four device groups, one trigger, one episode file. The clock node is the only part of this diagram that is hard.
Measured sync error
Absolute timestamp offset between the head camera and every other stream, sampled across randomly selected frames from shipped episodes.
Method, tooling and the raw CSV behind this plot are in the sync measurement guide. Nobody else in this category publishes this number, which is exactly why we do.

Coverage — and what we don't have
Everyone publishes a bigger number. We publish the histogram, including the gaps.
If your task lives in a thin column, that is a spec conversation, not a no. Thin cells are where we are actively accepting specs, and where a pilot buys you exclusivity on the capture.
Evidence it trains better
The only number that matters is the one downstream of the data.
Same architecture, same eval suite, same seeds — only the pretraining corpus changes. Hours are on the label so you can see we are not winning on volume.
Training code, eval harness and per-task breakdown live in the benchmark repo. If a number here does not reproduce on your cluster, we want the issue filed.
- Architecture
- Identical diffusion policy, 210 M params, no per-dataset tuning
- Eval suite
- 40 unseen manipulation tasks, 20 trials each, fixed seeds
- Success criterion
- Task-completion predicate, scored by held-out rubric, not by us
- Confound
- Firsthand is 1,200 h vs EgoScale 20,854 h — the delta is fitness, not volume
Provenance and consent
The license is a product surface, not a footnote.
Your legal team will read this section before your research team reads the specs. So here is the actual release language, the actual pay, and the actual de-identification pipeline.
Participant release — verbatim
What contributors are paid
- Base session rate
- $38 / hour
- Includes
- Setup, calibration, breaks, travel over 20 min
- Thin-coverage premium
- +$12 / hour (clinical, construction, agriculture)
- Payment terms
- Net 7, no per-task bounty, no rejection clawback
Published because a per-task bounty produces rushed, unusable footage. Collector standards.
De-identification pipeline
- Faces
- Bystander faces detected and blurred; participant faces mostly out of frame by rig geometry
- Plates
- Licence plates and street numbers blurred on all outdoor and automotive footage
- Audio
- Speech scrubbed of names, addresses and card numbers; ambient audio retained
- Screens
- Visible phone and monitor content masked unless it is the manipulated object
- Bystanders
- If a person cannot be de-identified, the clip is discarded rather than shipped
Jurisdictions of capture
- United States
- Canada
- United Kingdom
- Germany
- Netherlands
- Poland
- Portugal
- Japan
- South Korea
- Singapore
- Mexico
- Brazil
- GDPR ART. 6(1)(a)
- CCPA / CPRA
- UK GDPR
- APPI (JP)
- PIPEDA (CA)
- SOC 2 TYPE II — IN AUDIT
Case study
A bimanual loading policy that kept failing on half-full racks.
Anonymised at the buyer's request: a US humanoid lab, Series B, training a bimanual dishwasher-loading policy. Their footage was clean and their policy still failed the moment a rack was partially occupied.
- 620hours
- Validated hours delivered in 11 weeks
- 11skills
- Named skills in the spec, all cells filled
- +21pts
- Success rate on their internal dishwasher eval
- 7.4%
- Batch reject rate — recaptured at our cost
The lab had 4,000 hours of public egocentric footage and a policy that scored well in simulation. In the real world it stalled on any rack that was already half loaded, because almost nothing in the public corpora shows a person recovering from a bad grasp inside a cluttered fixture.
We wrote a spec with a hard quota: 15% of every batch had to be a failure and recovery — dropped plate, wrong slot, tine collision — with the recovery captured through to completion. That quota is the whole case study.
Their words, lightly trimmed: “The failure cases were the only part we couldn’t buy anywhere else, and they moved the eval more than the other 500 hours combined.”

Engagement timeline
- WEEK 0Spec written together: 11 skills, 6 kitchens, explicit failure-case quota of 15%
- WEEK 1Rig calibration on their object set — 34 SKUs of real crockery shipped to us
- WEEK 2First 50 h batch delivered in RLDS. They found a wrist extrinsics error. We recaptured 12 h.
- WEEK 4Cadence at 70 h/week. Coverage report per batch with the reject log attached.
- WEEK 11620 h delivered. Failure cases turned out to be the highest-value slice.
Anonymised with permission. Reference call available under NDA once a spec is scoped.
Pipeline
Five steps. Spec to delivery.
Step 03 is the one most vendors leave out, and it is the reason you are not paying for footage you will throw away.
- 01SPECNamed skill list, coverage targets, accept/reject criteria
- 02CAPTURECalibrated rigs, vetted and trained operators, real environments
- 03VALIDATESync check, QA pass, reject and recapture — you see the reject log
- 04ENRICHDepth, body pose, hand pose, action segments, instance masks
- 05DELIVERRLDS, LeRobot, WebDataset, HDF5 — plus the data card
FAQ
Questions we get on the first call.
Answered here so the first call can be about your skill spec instead.
01What is egocentric video data?
Egocentric video is footage recorded from a head-mounted camera worn by the person performing a task, so the camera shares the actor’s viewpoint and moves with their head. For embodied AI it matters because the observation distribution matches what a robot with a head or chest camera actually sees: hands entering frame from below, occlusion during grasp, motion blur during reach, and a first-person view of the manipulated object. Third-person footage does not contain that geometry, so policies trained on it transfer poorly.
02How much data do I need to train an embodied AI model?
NVIDIA’s EgoScale work reports a log-linear relationship between egocentric hours and downstream policy success out to 20,854 hours with no observed saturation, so more hours keep helping. In practice the binding constraint is fitness, not volume: teams typically discard around 90% of bulk footage as unusable for manipulation. A useful starting point is 50–100 validated hours against one named skill, which is enough to measure whether your architecture responds to the data before you commit to a thousand hours.
03How is this different from Ego4D or EgoDex?
Ego4D’s license does permit commercial use, so the honest gap is fitness rather than legality: it is roughly 3,670 hours of unstructured daily activity, much of it with heavy motion blur, no manipulation-grade 3D hand pose, and no frame-accurate action labels. EgoDex has excellent hand pose but is licensed CC BY-NC-ND, which genuinely rules out commercial training use. Firsthand captures against a named skill spec, hardware-synchronizes every stream to under two milliseconds, ships depth, 3D hand pose, body pose, segmentation and action segments on every clip, and attaches a commercial buyer-owned license with a documented consent chain.
04What hardware do you capture on?
A head-mounted global-shutter 4K/60 camera with a 120° lens, two wrist-mounted 1080p/60 cameras, an active-stereo depth sensor at 848×480/30, and 200 Hz 6-DoF IMUs. All streams are hardware-triggered onto a single PTP clock domain, and intrinsics and extrinsics are re-solved with a checkerboard target at the start and end of every session so drift is detectable rather than assumed.
05What does it cost?
We quote per engagement rather than publishing a rate card, because the price is driven by things we cannot guess from a web page: which environments you need, how thin the coverage is in them, how dense the annotation has to be, and whether you want exclusivity. The sample pack is free and ungated, so you can evaluate the data before any commercial conversation. Tell us what you are training and your target hours in the intake form and you will get a scoped estimate — hours, price, delivery date, and the coverage cells we would need to fill — within one business day. Billing is always per validated hour: footage that passed sync and QA against your spec. Rejected episodes, recaptures, calibration time and travel are never billed.
06How do you handle consent and licensing?
Every participant signs a written release granting commercial training rights before capture begins, and is paid a published hourly rate rather than a per-task bounty. Faces and licence plates of bystanders are blurred, audio is scrubbed of identifying speech, and any clip where a bystander cannot be de-identified is discarded rather than shipped. You receive a perpetual, irrevocable, buyer-owned commercial license, the consent artefacts for the batch, and a data card listing jurisdictions of capture.
07What formats do you deliver in?
RLDS, LeRobot, WebDataset, HDF5, zarr, and Rerun .rrd. Every episode carries int64 nanosecond timestamps per stream, camera intrinsics and extrinsics, handJoints[2][21][3], bodyJoints[24][3], instance masks, and action segments. Converters ship as readable source, not a binary, so you can retarget the schema to your own loader.
08How fast can you start?
Spec conversation the same week, first validated batch in a median of nine days from a signed spec. Thin coverage areas — construction, clinical, agriculture, hospitality — take longer to staff, typically three to four weeks to first batch, because operator vetting in those environments is the bottleneck.
Request a quote
Tell us what you're training.
We quote per engagement rather than off a rate card, because price depends on which environments you need, how thin our coverage is in them, and how dense the annotation has to be. Fill this in and you get real numbers back.
- Reply time
- One business day, with a scoped estimate
- You receive
- Hours, price, delivery date, coverage cells to fill
- First call
- 30 minutes, technical, no deck
- Then
- A written spec you own, whether or not you buy
- Billing unit
- Validated hours only — rejects are never billed
Or skip the form entirely — hello@ifirsthand.com. A research engineer reads it, not a sales inbox.