Solutions / Multimodal & sensor
Multiple streams, captured on one clock.
Synchronized combinations of video, audio, depth, motion, and pose — for robotics, embodied AI, and multimodal models that need the streams to line up, not just coexist in the same folder.

Definition
Multimodal data collection
Multimodal data collection is the simultaneous capture of two or more aligned data streams — for example video, depth, audio, and motion — from the same event, time-synchronized so a model can learn the relationships between them.
What we collect
The formats teams ask us for.
Not an exhaustive menu — if what you need is not here, it is a custom collection, which we also do.
Synchronized video + depthRGB and metric depth aligned per frame for 3D perception and manipulation.
Motion & IMUAccelerometer and gyroscope streams, hardware-timestamped on the same clock as video.
Hand & body pose3D hand and body keypoints tracked through the sequence, with visibility and contact events.
Audio + visualAligned speech or event audio with video for audio-visual models.
Sensor & robotics dataMulti-sensor rigs for robotics and IoT, aligned to a single timeline.
Egocentric multi-streamOur core rig: head and wrist cameras, depth, IMU, and pose synchronized under two milliseconds.
What shapes the spec
The decisions we settle before collecting.
- Synchronization
- Streams hardware-triggered onto a shared clock; egocentric rig verified under 2 ms
- Stream set
- Chosen per project — video, depth, audio, IMU, pose, segmentation
- Calibration
- Intrinsics and extrinsics re-solved per session so drift is detectable, not assumed
- Delivery
- RLDS, LeRobot, WebDataset, HDF5, zarr, or Rerun .rrd, with per-stream timestamps
- Consent basis
- Written release per participant; discard-don’t-downgrade on de-identification
Where it is used
What teams train with it.
- Robot manipulation and imitation learning
- Embodied AI and world models
- Audio-visual understanding
- Activity recognition from fused sensors
- IoT and wearable sensor models
How we run it
The standards behind every batch.
These apply to every modality — they are the reason the data is usable rather than merely large.
- Collected to a written spec
- Nothing is gathered speculatively. We agree the target — languages, demographics, devices, environments, edge cases — before a single contributor is briefed.
- Contributors paid hourly
- People are paid for their time, including setup and retakes, not a bounty per item. A per-item rate optimises for volume and quietly wrecks quality.
- Documented, informed consent
- Every contributor signs a release granting the usage rights you need before collection begins. You receive the consent artefacts with the batch.
- Discard rather than downgrade
- If an item cannot meet the spec or a bystander cannot be de-identified, it is dropped — not shipped at a discount to pad the count.
- Multi-layer QA with a visible reject log
- Automated checks plus human review, and you see the reject reasons, not just the accepted items.
- Buyer-owned commercial license
- You receive a perpetual, buyer-owned license with a data card recording jurisdictions of capture and the consent basis.
FAQ
Multimodal & sensor collection, answered.
- What does multimodal data collection actually involve?
- Capturing two or more streams from the same event — say video, depth, audio, and motion — and aligning them on a single clock so a model can learn how they relate. The hard part is synchronization: streams that are merely in the same folder but drift against each other are close to useless for manipulation. We hardware-trigger onto a shared clock and verify the offset.
- How tightly are the streams synchronized?
- For our egocentric rig, every stream is hardware-triggered onto one PTP clock domain and the median cross-stream offset is verified under two milliseconds. For other sensor combinations we agree and measure the synchronization target in the spec, and report the achieved offset with the batch.
- What formats do you deliver multimodal data in?
- RLDS, LeRobot, WebDataset, HDF5, zarr, and Rerun .rrd, with int64 nanosecond timestamps per stream and calibration included. Converters ship as readable source so you can retarget to your own loader.
Scope a multimodal & sensor collection.
Bring your spec or your problem. You will get a scoped estimate — reach, timeline, and price — before any commitment.