Shipping · 2,840 validated hours
Egocentric kitchen video dataset
- Episodes
- 18,420
- Participants
- 214
- Capture sites
- 96
- Sync ceiling
- 2.0 ms, hardware-triggered
What a frame contains
Every layer, on every clip in this set.
Rendered from the delivered schema rather than a marketing composite. Toggle layers to see what arrives with the footage.
Synthetic frame, real schema. The sample pack contains the same fields for actual episodes.

Named skills
What the operators were told to do.
Coverage is specified per skill, not per hour. These are the skill lines currently filled in this environment.
- Knife work — dice, julienne, chiffonade
- Pour and decant between vessels
- Open jars, bottles, vacuum-sealed packaging
- Load and unload dishwasher
- Stovetop transfer with a loaded pan
- Wipe and reset a work surface
- Bimanual dough handling
Environments captured
- Domestic kitchen
- Shared-house kitchen
- Small commercial prep line
Condition cells filled
- Daylight
- 94% of target cells filled
- Low light
- 61% — evening capture is the active gap
- Cluttered
- 79% — counters left as found, not staged
- Multi-person
- 44% — two operators sharing one counter
- Failure cases
- 28% — drops, spills, mis-grasps, recoveries
Known failure modes
What goes wrong in this environment.
Published because you will find these in the data within an afternoon, and it is cheaper for both of us if you find them in this list first.
- Wet and reflective surfaces degrade active-stereo depth; per-frame confidence maps are shipped so you can mask rather than guess.
- Steam events are labelled as an episode-level flag because they wreck depth for 2–6 seconds at a time.
What ships with every clip
Ten streams, one clock, one episode file.
Not an à la carte menu. Every validated hour we deliver carries the whole stack, in the schema below, whether you asked for depth or not.
| Stream | Spec | Detail |
|---|---|---|
| Head camera | 3840 × 2160 · 60 fps | Global-shutter, 120° HFOV, rolling-shutter-free, H.265 + lossless keyframes |
| Wrist cameras × 2 | 1920 × 1080 · 60 fps | Left + right, 100° HFOV, rigid mount, extrinsics re-solved per session |
| Depth | 848 × 480 · 30 fps | Active stereo, 0.3–4 m range, metric millimetres, per-frame confidence map |
| IMU | 200 Hz · 6-DoF | Accel + gyro, bias-calibrated, hardware-timestamped on the same clock domain |
| Hand pose | 21 keypoints × 2 hands | 3D metric, per-joint visibility flag, contact events on grasp and release |
| Body pose | 24 joints | 3D, root-relative and world-frame, torso and forearm chains resolved |
| Segmentation | Instance masks | Manipulated objects + target surfaces, tracked IDs across the episode |
| Action segments | Verb + noun taxonomy | 97 verbs, 512 nouns, start/end to the frame, human-reviewed |
| Formats | RLDS · LeRobot · WebDataset | Also HDF5, zarr, and .rrd for Rerun. Converters shipped as source. |
| License | Commercial · buyer-owned | Perpetual, irrevocable, model-weights-clean. Exclusivity available. |
Colour dots map to the modality legend used in every chart on this site. Full field-level schema in the episode schema docs.
Data card
The card that ships with the batch.
Delivered as machine-readable JSON alongside the episodes, so provenance travels with the data instead of living in an email thread.
- Dataset
- Egocentric kitchen video dataset
- Validated hours
- 2,840 h — passed sync and QA, shipped to at least one buyer
- Episodes / participants / sites
- 18,420 / 214 / 96
- Sync ceiling
- Δ < 2.0 ms across all streams · median 1.12 ms · p99 1.94 ms
- Consent
- Written commercial release per participant, signed before capture begins
- De-identification
- Faces and licence plates blurred, audio scrubbed, un-blurrable clips discarded
- Collection period
- Rolling. Batch capture dates recorded per episode.
- License
- Perpetual, irrevocable, buyer-owned commercial. Exclusivity available.
- Known limitations
- Wet and reflective surfaces degrade active-stereo depth; per-frame confidence maps are shipped so you can mask rather than guess.
- Formats
- RLDS, LeRobot, WebDataset, HDF5, zarr, Rerun .rrd
Flagged frames are shipped rather than silently dropped. You decide whether to mask, down-weight, or exclude them.
Nearest public dataset
How this compares to EPIC-KITCHENS-100.
Where a public set is genuinely better at something, we say so. Where the blocker is the license rather than the quality, we say that too.
EPIC-KITCHENS-100
- Hours
- 100 h
- License
- CC BY-NC 4.0 — non-commercial
Excellent action labels and the reference taxonomy for this environment, but no metric depth, no 3D hand pose, and a non-commercial license. We ship the enrichment layers it lacks under a license you can train on.
Firsthand — Kitchen & food prep
- Hours
- 2,840 h
- License
- Commercial, buyer-owned, perpetual
Captured against the named skill list above, hardware-synchronized under 2.0 ms, and shipped with depth, 3D hand pose, body pose, instance masks and action segments on every clip.
See the downstream result →Pilot kitchen & food prep against your spec.
Write the skill list with us, get fifty to a hundred validated hours in two weeks, and read the reject log before you commit to volume.