# Firsthand > Firsthand collects custom egocentric (first-person) video datasets for embodied AI and robotics: head + dual wrist cameras, depth, IMU, 3D hand and body pose, instance segmentation and action segments, hardware-synchronized under 2 ms (median 1.12 ms), with documented participant consent and a perpetual, buyer-owned commercial license. Billed per validated hour. Delivered in RLDS, LeRobot, WebDataset, HDF5, zarr and Rerun .rrd. ## Key facts - 10,585 validated hours; 16 countries; median 48 h to first delivery in deep-coverage environments - Contributors paid $38/h base (+$12/h thin-coverage premium) - Free sample pack: 40 episodes, 4 environments, 2.4 GB - Cross-stream sync: median 1.12 ms, p99 1.94 ms - Commercial-use license (contrast with Ego4D terms and EgoDex CC BY-NC-ND) ## Pages - [Egocentric Video Data for Embodied AI & Robotics](https://www.ifirsthand.com): Custom egocentric video datasets for embodied AI and robotics: ten synchronized streams, documented consent, and a buyer-owned commercial license. Billed per validated hour. - [Dataset catalog — coverage by environment](https://www.ifirsthand.com/data): Every environment we hold coverage in, with validated hours, condition-cell fill rates, and the honest gaps. - [Example clips — raw egocentric footage](https://www.ifirsthand.com/examples): Every raw egocentric clip Firsthand publishes, straight off the camera, downloadable without a form. - [Custom data collection across every modality](https://www.ifirsthand.com/solutions): Custom data collection for AI training across audio, image, video, text, and multimodal — plus fully custom worldwide collection. - [Custom worldwide data collection](https://www.ifirsthand.com/solutions/custom-collection): When the training data you need does not exist yet, we build the collection program around it — anywhere in the world. - [Embodied AI training data](https://www.ifirsthand.com/embodied-ai-data): Synchronized egocentric video, depth, motion, and pose for embodied AI, world models, and robot learning — collected to spec with documented consent. - [Robot manipulation data](https://www.ifirsthand.com/robot-manipulation-data): Human-demonstration data for robot manipulation and imitation learning: egocentric video, hand pose, and contact events in RLDS and LeRobot. - [Capture methodology](https://www.ifirsthand.com/capture): How Firsthand captures egocentric video: hardware-triggered multi-sensor rigs on a single PTP clock, per-session calibration, and a QA pass with a published reject log. - [Automatic anonymization](https://www.ifirsthand.com/anonymization): Faces, plates, and screens detected and irreversibly blurred at ingest, confirmed by human review, before any footage is delivered. - [Documentation — episode schema & formats](https://www.ifirsthand.com/docs): Field-level episode schema, RLDS / LeRobot / WebDataset / HDF5 / zarr / Rerun delivery formats, and a Python quickstart. - [Evaluation methodology & benchmarks](https://www.ifirsthand.com/benchmarks): How we evaluate egocentric pretraining data: a fixed policy architecture, a held-out manipulation suite, and honest comparisons. - [Provenance — consent, pay & licensing](https://www.ifirsthand.com/provenance): The participant release verbatim, contributor pay rates, the de-identification pipeline, and how commercial licensing compares to public datasets. - [Company](https://www.ifirsthand.com/company): Firsthand builds egocentric capture programs for embodied AI teams: how we are structured, what we refuse to do, and how to reach us. - [For capture operators](https://www.ifirsthand.com/for-collectors): Standards, pay rates, and the acceptance bar for capture operators and partner collection teams contributing egocentric hours to Firsthand. - [Audio & speech data collection](https://www.ifirsthand.com/solutions/audio-speech): Custom speech and audio data collection for ASR, voice assistants, and audio models: scripted and spontaneous speech, multilingual and accented, recorded to your spec with documented consent. - [Image data collection](https://www.ifirsthand.com/solutions/image): Custom image data collection for computer vision: real-world photos across devices, lighting, and geographies, sourced from a vetted global crowd with documented consent and to your spec. - [Video data collection](https://www.ifirsthand.com/solutions/video): Custom video data collection for action recognition, gesture, and activity understanding — including egocentric first-person capture — sourced worldwide with documented consent. - [Text & language data collection](https://www.ifirsthand.com/solutions/text-language): Custom text and language data collection: prompts, conversations, translations, and domain text authored by native speakers worldwide for LLM training, evaluation, and fine-tuning. - [Multimodal & sensor data collection](https://www.ifirsthand.com/solutions/multimodal): Custom multimodal and sensor data collection: synchronized video, audio, depth, IMU, and pose for robotics, embodied AI, and multimodal models — captured to spec with documented consent. - [Anonymization as a service — face, plate & video redaction at scale](https://www.ifirsthand.com/services): Send us footage, get it back de-identified: faces, plates, screens and speech automatically detected and irreversibly blurred, confirmed by human review, with a per-file certificate. - [Face blurring service — blur faces in video at scale](https://www.ifirsthand.com/services/face-blurring): Automatically detect and irreversibly blur faces in video and images at scale. A vision model returns a bounding box and confidence per face; each region is blurred into the delivered pixels and confirmed by human review. - [Video anonymization service — faces, plates, screens & speech](https://www.ifirsthand.com/services/video-anonymization): A full de-identification pass for video: faces, licence plates, legible screens and documents detected and irreversibly blurred, identifying speech scrubbed from audio, and every clip confirmed by human review with a per-file certificate. - [Licence plate blurring & redaction service for video](https://www.ifirsthand.com/services/license-plate-blurring): Automatically detect and irreversibly blur licence plates in video and images at scale — dashcam, mapping, car-park, and CCTV footage. Plates tracked across frames, redacted into the pixels, and confirmed by human review. - [Kitchen & food prep](https://www.ifirsthand.com/datasets/egocentric-kitchen-video): Egocentric kitchen video dataset: 2,840 validated hours of two-handed food prep in real homes, with depth, 3D hand pose, body pose, instance masks and action segments. Commercial license, documented consent. - [Warehouse & logistics](https://www.ifirsthand.com/datasets/warehouse-picking): Warehouse picking egocentric dataset: 1,420 validated hours of pick, scan, pack and pallet-build captured on shift in live fulfilment space, with depth, hand pose and action segments under a commercial license. - [Retail & shelf ops](https://www.ifirsthand.com/datasets/retail-shelf-stocking): Retail shelf stocking egocentric dataset: 890 validated hours of facing, restocking, price changes and planogram resets on live retail floors, fully de-identified, with depth and 3D hand pose. - [Automotive service](https://www.ifirsthand.com/datasets/automotive-service-bay): Automotive service bay egocentric dataset: 640 validated hours of torque, fluid, trim and diagnostic work in working bays, with depth-quality flags, 3D hand pose and action segments. Plates blurred at ingest. - [Tool use & assembly](https://www.ifirsthand.com/datasets/tool-use-assembly): Tool use and assembly egocentric dataset: 1,680 validated hours of bench assembly, fastening, measuring and fitting with 3D hand pose, contact events and frame-accurate action segments. - [Household chores](https://www.ifirsthand.com/datasets/household-chores): Household chores egocentric dataset: 2,310 validated hours of laundry, tidying, cleaning and surface reset across whole homes, with per-frame instance masks for deformable objects and 3D hand pose. - [Compare egocentric video data providers & datasets](https://www.ifirsthand.com/compare): Firsthand compared with Ego4D, EgoDex and EgoScale, a public dataset license matrix, and a neutral vendor checklist. - [Firsthand vs Ego4D — commercial egocentric data compared](https://www.ifirsthand.com/compare/ego4d): Ego4D vs Firsthand for embodied AI: hours, hand pose, action labels, sync, license and commercial use compared. When each is the right training corpus. - [Firsthand vs EgoDex — hand pose quality and license compared](https://www.ifirsthand.com/compare/egodex): EgoDex vs Firsthand for robot learning: EgoDex has strong 3D hand pose but a CC BY-NC-ND non-commercial license. Compare streams, sync, consent and commercial use. - [Firsthand vs EgoScale — volume vs fitness for robot policies](https://www.ifirsthand.com/compare/egoscale): EgoScale vs Firsthand: EgoScale shows log-linear scaling to ~20,854 hours; Firsthand optimizes fitness per validated hour. Compare streams, labels, license and cost. - [Public egocentric dataset licenses — can you train commercially?](https://www.ifirsthand.com/compare/public-dataset-licenses): A plain-English license comparison of public egocentric datasets (Ego4D, EgoDex, EgoScale, EPIC-Kitchens) versus Firsthand: what each allows for commercial robot training. - [How to choose an egocentric video data vendor](https://www.ifirsthand.com/compare/egocentric-data-providers): A neutral checklist for evaluating egocentric video data providers: sync error, annotation depth, license, consent chain, reject policy, billing unit and time to first batch. - [Egocentric AI & embodied data glossary](https://www.ifirsthand.com/glossary): Plain-language definitions of the terms used in egocentric video and embodied-AI training data. - [Egocentric video — definition](https://www.ifirsthand.com/glossary/egocentric-video): First-person video recorded from a camera worn by the person doing the task. - [Embodied AI — definition](https://www.ifirsthand.com/glossary/embodied-ai): AI that perceives and acts through a physical or simulated body. - [Skill spec — definition](https://www.ifirsthand.com/glossary/skill-spec): A written definition of the exact task, environments and conditions to capture. - [Cross-stream sync error — definition](https://www.ifirsthand.com/glossary/cross-stream-sync-error): The timing offset between simultaneously captured streams like video, depth and IMU. - [Action segment — definition](https://www.ifirsthand.com/glossary/action-segment): A labeled time interval of one sub-action, with start/end timestamps and a verb-noun. - [Validated hour — definition](https://www.ifirsthand.com/glossary/validated-hour): An hour of footage that has passed QA against the spec — not an hour of raw recording. - [Consent chain — definition](https://www.ifirsthand.com/glossary/consent-chain): The documented trail proving every participant granted the rights you need. - [Hand pose estimation — definition](https://www.ifirsthand.com/glossary/hand-pose-estimation): Recovering 3D positions of hand and finger joints over time. - [Failure-case coverage — definition](https://www.ifirsthand.com/glossary/failure-case-coverage): Deliberately including the hard cases: transit, occlusion, low light, recovery. - [RLDS — definition](https://www.ifirsthand.com/glossary/rlds): Reinforcement Learning Datasets — an episodic format for robot learning data. - [LeRobot format — definition](https://www.ifirsthand.com/glossary/lerobot-format): An episodic dataset format used by the LeRobot robot-learning ecosystem. - [WebDataset — definition](https://www.ifirsthand.com/glossary/webdataset): A tar-based format for streaming large datasets efficiently during training. - [Imitation learning — definition](https://www.ifirsthand.com/glossary/imitation-learning): Training a policy to reproduce demonstrated behavior from example trajectories. - [Teleoperation vs egocentric data — definition](https://www.ifirsthand.com/glossary/teleoperation-vs-egocentric-data): Robot-collected demonstrations versus human-worn first-person capture. - [Global-shutter camera — definition](https://www.ifirsthand.com/glossary/global-shutter-camera): A camera that exposes every pixel simultaneously, avoiding rolling-shutter skew. - [PTP clock sync — definition](https://www.ifirsthand.com/glossary/ptp-clock-sync): Precision Time Protocol — sub-microsecond clock alignment across sensors. - [Instance segmentation — definition](https://www.ifirsthand.com/glossary/instance-segmentation): Per-pixel masks that identify and separate individual objects over time. - [Data card — definition](https://www.ifirsthand.com/glossary/data-card): A structured record of how a dataset was collected, consented and licensed. - [De-identification — definition](https://www.ifirsthand.com/glossary/de-identification): Irreversibly removing personal identifiers like faces, plates and screens. - [Bimanual manipulation — definition](https://www.ifirsthand.com/glossary/bimanual-manipulation): Tasks using both hands together, often with coordinated, asymmetric roles. - [Face detection — definition](https://www.ifirsthand.com/glossary/face-detection): Locating faces in an image and returning a box and confidence for each. - [Bounding box — definition](https://www.ifirsthand.com/glossary/bounding-box): The rectangle of pixel coordinates that marks where a detected object is. - [Anonymization — definition](https://www.ifirsthand.com/glossary/anonymization): Making a person no longer identifiable from data by any reasonably likely means. - [Redaction — definition](https://www.ifirsthand.com/glossary/redaction): Permanently obscuring specific regions of an image, video, or document. - [Guides — egocentric data collection & training](https://www.ifirsthand.com/guides): Practical guides for collecting and training on egocentric data: skill specs, sync measurement, calibration, evaluation, and data volume. - [How to write a skill spec | Firsthand](https://www.ifirsthand.com/guides/how-to-write-a-skill-spec): A skill spec names the task, geometry, condition budget, reject rule and one acceptance test. Template and worked example for egocentric data collection. - [How we measure sync error | Firsthand](https://www.ifirsthand.com/guides/how-we-measure-sync-error): Cross-stream sync error is the residual timing offset between RGB, depth, IMU and pose. How Firsthand hardware-triggers and measures it: median 1.12 ms, p99 1.94 ms. - [Camera calibration procedure | Firsthand](https://www.ifirsthand.com/guides/camera-calibration-procedure): How Firsthand calibrates a head + dual-wrist egocentric rig: opening and closing solves, RMS reprojection under 0.28 px, and a drift tolerance measured per episode. - [Evaluation methodology | Firsthand](https://www.ifirsthand.com/guides/evaluation-methodology): How to evaluate whether an egocentric dataset trains a better manipulation policy: held-out environments, matched hour budgets, and the open ego-eval protocol. - [Python quickstart | Firsthand](https://www.ifirsthand.com/guides/python-quickstart): Install the firsthand-ego package, list your access, pull an episode in RLDS or WebDataset, and stream aligned RGB, depth and hand pose into a PyTorch loop. - [How much egocentric data do you need? | Firsthand](https://www.ifirsthand.com/guides/how-much-egocentric-data-do-you-need): A practical starting point is 50–100 validated hours per skill, then scale where evaluation is thin. Why validated hours, not raw hours, are the unit that matters. - [Egocentric data for humanoid robots | Firsthand](https://www.ifirsthand.com/guides/egocentric-data-for-humanoid-robots): Humanoids with head and wrist cameras need first-person, bimanual demonstration data with 3D hand pose. What to capture and how it maps to a humanoid embodiment. - [Licensing training data for commercial models | Firsthand](https://www.ifirsthand.com/guides/how-to-license-training-data-for-commercial-models): Commercial model teams need documented consent and a perpetual, buyer-owned commercial license — not a research-only dataset. What to check before you train. - [How to blur faces in video automatically | Firsthand](https://www.ifirsthand.com/guides/how-to-blur-faces-in-video-automatically): Automatically blur faces in video with a five-step computer-vision pipeline: detect each face as a bounding box, track it across frames, blur it into the pixels, review, and certify. Method and pitfalls. - [Face detection & bounding boxes explained | Firsthand](https://www.ifirsthand.com/guides/face-detection-bounding-boxes-explained): What a face-detection model returns — bounding-box coordinates and a confidence score — how the threshold trades recall against precision, and why detection is the first step of anonymization, not the whole of it. - [Video anonymization & privacy law (GDPR, CCPA) | Firsthand](https://www.ifirsthand.com/guides/video-anonymization-and-privacy-law): How video anonymization relates to GDPR and CCPA: what counts as personal data in footage, the line between pseudonymization and anonymization, why redaction must be irreversible, and what an auditable record needs to show. - [Delivery formats — RLDS, LeRobot, WebDataset](https://www.ifirsthand.com/formats): Every Firsthand delivery format explained: RLDS, LeRobot, WebDataset, HDF5, zarr and Rerun .rrd, plus the canonical episode schema. - [RLDS delivery format | Firsthand](https://www.ifirsthand.com/formats/rlds): How Firsthand delivers egocentric episodes in RLDS: TFDS-compatible shards, step and episode keys, and a readable converter you can retarget to your own loader. - [LeRobot delivery format | Firsthand](https://www.ifirsthand.com/formats/lerobot): How Firsthand delivers egocentric episodes in the LeRobot v2.1 layout: parquet plus mp4 per camera, an observation.state vector, and the smallest on-disk option. - [WebDataset delivery format | Firsthand](https://www.ifirsthand.com/formats/webdataset): How Firsthand delivers egocentric episodes as WebDataset tar shards: 512 samples per shard, signed-URL streaming, and per-sample keys for RGB, depth and annotations. - [HDF5 & zarr delivery format | Firsthand](https://www.ifirsthand.com/formats/hdf5-zarr): How Firsthand delivers egocentric episodes as HDF5 or zarr: one group per episode, time-major chunking, and per-stream nanosecond timestamps for random access and sync checks. - [Rerun .rrd delivery format | Firsthand](https://www.ifirsthand.com/formats/rerun-rrd): How Firsthand delivers egocentric episodes as Rerun .rrd recordings: every stream on one timeline, with camera, hand, point-cloud and sync-error entities to inspect first. - [Episode schema | Firsthand](https://www.ifirsthand.com/formats/episode-schema): The canonical Firsthand episode layout: meta, streams, annotations and calibration groups, with dtypes, shapes and units. Every delivery format is a projection of this. - [Terms of Service | Firsthand](https://www.ifirsthand.com/terms): The terms governing use of the Firsthand website, sample datasets, and data-collection services. Placeholder pending legal review. - [Privacy Policy | Firsthand](https://www.ifirsthand.com/privacy): How Firsthand handles website visitor data and the consented participant footage in its datasets. Placeholder pending legal review. - [Data Processing Addendum | Firsthand](https://www.ifirsthand.com/dpa): The data-processing terms that attach to Firsthand commercial agreements, covering roles, security, and sub-processors. Placeholder pending legal review. ## Contact - Sample pack and quotes: https://www.ifirsthand.com/#contact