Shipping · 1,420 validated hours
Warehouse picking egocentric dataset
- Episodes
- 9,060
- Participants
- 138
- Capture sites
- 22
- Sync ceiling
- 2.0 ms, hardware-triggered
What a frame contains
Every layer, on every clip in this set.
Rendered from the delivered schema rather than a marketing composite. Toggle layers to see what arrives with the footage.
Synthetic frame, real schema. The sample pack contains the same fields for actual episodes.

Named skills
What the operators were told to do.
Coverage is specified per skill, not per hour. These are the skill lines currently filled in this environment.
- Single-item pick from bin to tote
- Barcode scan with a handheld gun
- Carton break-down and flattening
- Pallet stacking and shrink-wrap
- Label application and print reconciliation
- Two-person long-item carry
Environments captured
- Pick aisle
- Pack bench
- Pallet build
- Returns desk
Condition cells filled
- Mixed light
- 76% — high-bay sodium plus daylight doors
- Low light
- 38% — night shift capture underway
- Cluttered
- 66%
- Multi-person
- 58% — the highest of any environment
- Failure cases
- 19% — mis-picks, dropped totes, jams
Known failure modes
What goes wrong in this environment.
Published because you will find these in the data within an afternoon, and it is cheaper for both of us if you find them in this list first.
- Long sight-lines push depth beyond the 4 m reliable range; distance-invalid pixels are flagged, not interpolated.
- Rig cabling is routed under hi-vis PPE at every site, which changes wrist extrinsics slightly — hence per-session recalibration.
What ships with every clip
Ten streams, one clock, one episode file.
Not an à la carte menu. Every validated hour we deliver carries the whole stack, in the schema below, whether you asked for depth or not.
| Stream | Spec | Detail |
|---|---|---|
| Head camera | 3840 × 2160 · 60 fps | Global-shutter, 120° HFOV, rolling-shutter-free, H.265 + lossless keyframes |
| Wrist cameras × 2 | 1920 × 1080 · 60 fps | Left + right, 100° HFOV, rigid mount, extrinsics re-solved per session |
| Depth | 848 × 480 · 30 fps | Active stereo, 0.3–4 m range, metric millimetres, per-frame confidence map |
| IMU | 200 Hz · 6-DoF | Accel + gyro, bias-calibrated, hardware-timestamped on the same clock domain |
| Hand pose | 21 keypoints × 2 hands | 3D metric, per-joint visibility flag, contact events on grasp and release |
| Body pose | 24 joints | 3D, root-relative and world-frame, torso and forearm chains resolved |
| Segmentation | Instance masks | Manipulated objects + target surfaces, tracked IDs across the episode |
| Action segments | Verb + noun taxonomy | 97 verbs, 512 nouns, start/end to the frame, human-reviewed |
| Formats | RLDS · LeRobot · WebDataset | Also HDF5, zarr, and .rrd for Rerun. Converters shipped as source. |
| License | Commercial · buyer-owned | Perpetual, irrevocable, model-weights-clean. Exclusivity available. |
Colour dots map to the modality legend used in every chart on this site. Full field-level schema in the episode schema docs.
Data card
The card that ships with the batch.
Delivered as machine-readable JSON alongside the episodes, so provenance travels with the data instead of living in an email thread.
- Dataset
- Warehouse picking egocentric dataset
- Validated hours
- 1,420 h — passed sync and QA, shipped to at least one buyer
- Episodes / participants / sites
- 9,060 / 138 / 22
- Sync ceiling
- Δ < 2.0 ms across all streams · median 1.12 ms · p99 1.94 ms
- Consent
- Written commercial release per participant, signed before capture begins
- De-identification
- Faces and licence plates blurred, audio scrubbed, un-blurrable clips discarded
- Collection period
- Rolling. Batch capture dates recorded per episode.
- License
- Perpetual, irrevocable, buyer-owned commercial. Exclusivity available.
- Known limitations
- Long sight-lines push depth beyond the 4 m reliable range; distance-invalid pixels are flagged, not interpolated.
- Formats
- RLDS, LeRobot, WebDataset, HDF5, zarr, Rerun .rrd
Flagged frames are shipped rather than silently dropped. You decide whether to mask, down-weight, or exclude them.
Nearest public dataset
How this compares to Ego4D (logistics subset).
Where a public set is genuinely better at something, we say so. Where the blocker is the license rather than the quality, we say that too.
Ego4D (logistics subset)
- Hours
- ~90 h
- License
- Commercial use permitted
Ego4D does allow commercial use, so the gap here is fitness: its logistics footage is incidental rather than spec-driven, unsynchronized across streams, and carries no manipulation-grade hand pose. Ours is captured against a named pick-and-pack skill list.
Firsthand — Warehouse & logistics
- Hours
- 1,420 h
- License
- Commercial, buyer-owned, perpetual
Captured against the named skill list above, hardware-synchronized under 2.0 ms, and shipped with depth, 3D hand pose, body pose, instance masks and action segments on every clip.
See the downstream result →Pilot warehouse & logistics against your spec.
Write the skill list with us, get fifty to a hundred validated hours in two weeks, and read the reject log before you commit to volume.