Shipping · 890 validated hours
Retail shelf stocking egocentric dataset
- Episodes
- 6,210
- Participants
- 96
- Capture sites
- 31
- Sync ceiling
- 2.0 ms, hardware-triggered
What a frame contains
Every layer, on every clip in this set.
Rendered from the delivered schema rather than a marketing composite. Toggle layers to see what arrives with the footage.
Synthetic frame, real schema. The sample pack contains the same fields for actual episodes.

Named skills
What the operators were told to do.
Coverage is specified per skill, not per hour. These are the skill lines currently filled in this environment.
- Face and front stock to a planogram
- Case-cut and shelf-load
- Price and label swap
- Chilled-case rotation by date code
- Hanger and fixture reset
Environments captured
- Grocery aisle
- Chilled case
- Apparel fixture
- Back-of-house staging
Condition cells filled
- Daylight
- 68%
- Mixed light
- 71%
- Multi-person
- 64% — customers in frame throughout
- Failure cases
- 14% — the thinnest failure coverage in the catalog
Known failure modes
What goes wrong in this environment.
Published because you will find these in the data within an afternoon, and it is cheaper for both of us if you find them in this list first.
- Every bystander is blurred at ingest; clips where a bystander cannot be de-identified are discarded, which is why episode yield here is lower than kitchen.
- Chilled-case capture is time-boxed to protect stock, so episodes are shorter — median 41 s versus 78 s elsewhere.
What ships with every clip
Ten streams, one clock, one episode file.
Not an à la carte menu. Every validated hour we deliver carries the whole stack, in the schema below, whether you asked for depth or not.
| Stream | Spec | Detail |
|---|---|---|
| Head camera | 3840 × 2160 · 60 fps | Global-shutter, 120° HFOV, rolling-shutter-free, H.265 + lossless keyframes |
| Wrist cameras × 2 | 1920 × 1080 · 60 fps | Left + right, 100° HFOV, rigid mount, extrinsics re-solved per session |
| Depth | 848 × 480 · 30 fps | Active stereo, 0.3–4 m range, metric millimetres, per-frame confidence map |
| IMU | 200 Hz · 6-DoF | Accel + gyro, bias-calibrated, hardware-timestamped on the same clock domain |
| Hand pose | 21 keypoints × 2 hands | 3D metric, per-joint visibility flag, contact events on grasp and release |
| Body pose | 24 joints | 3D, root-relative and world-frame, torso and forearm chains resolved |
| Segmentation | Instance masks | Manipulated objects + target surfaces, tracked IDs across the episode |
| Action segments | Verb + noun taxonomy | 97 verbs, 512 nouns, start/end to the frame, human-reviewed |
| Formats | RLDS · LeRobot · WebDataset | Also HDF5, zarr, and .rrd for Rerun. Converters shipped as source. |
| License | Commercial · buyer-owned | Perpetual, irrevocable, model-weights-clean. Exclusivity available. |
Colour dots map to the modality legend used in every chart on this site. Full field-level schema in the episode schema docs.
Data card
The card that ships with the batch.
Delivered as machine-readable JSON alongside the episodes, so provenance travels with the data instead of living in an email thread.
- Dataset
- Retail shelf stocking egocentric dataset
- Validated hours
- 890 h — passed sync and QA, shipped to at least one buyer
- Episodes / participants / sites
- 6,210 / 96 / 31
- Sync ceiling
- Δ < 2.0 ms across all streams · median 1.12 ms · p99 1.94 ms
- Consent
- Written commercial release per participant, signed before capture begins
- De-identification
- Faces and licence plates blurred, audio scrubbed, un-blurrable clips discarded
- Collection period
- Rolling. Batch capture dates recorded per episode.
- License
- Perpetual, irrevocable, buyer-owned commercial. Exclusivity available.
- Known limitations
- Every bystander is blurred at ingest; clips where a bystander cannot be de-identified are discarded, which is why episode yield here is lower than kitchen.
- Formats
- RLDS, LeRobot, WebDataset, HDF5, zarr, Rerun .rrd
Flagged frames are shipped rather than silently dropped. You decide whether to mask, down-weight, or exclude them.
Nearest public dataset
How this compares to MMAct / retail CCTV corpora.
Where a public set is genuinely better at something, we say so. Where the blocker is the license rather than the quality, we say that too.
MMAct / retail CCTV corpora
- Hours
- Varies
- License
- Research-only, mostly third-person
Existing retail footage is overwhelmingly fixed-camera surveillance, which gives you none of the first-person reach geometry. This is head- and wrist-mounted capture from the person doing the stocking.
Firsthand — Retail & shelf ops
- Hours
- 890 h
- License
- Commercial, buyer-owned, perpetual
Captured against the named skill list above, hardware-synchronized under 2.0 ms, and shipped with depth, 3D hand pose, body pose, instance masks and action segments on every clip.
See the downstream result →Pilot retail & shelf ops against your spec.
Write the skill list with us, get fifty to a hundred validated hours in two weeks, and read the reject log before you commit to volume.