Stereoscopic capture

Stereoscopic video data collection for 3D-aware AI.

Calibrated left- and right-eye video from a head-mounted stereo rig, hardware-synced with depth, motion, and hand pose. It is the binocular, human-baseline view that 3D perception and manipulation policies learn best from.
Views
Left + right, calibrated
Sync
Single hardware clock
Extras
Depth · IMU · hand pose
Formats
RLDS · HDF5 · zarr
A person wearing a head-mounted stereo camera works at a bench beside a monitor showing left and right views with a depth overlay.
Two lenses at a human interpupillary baseline, calibrated per session so depth is recoverable from every frame.

Definition

What is stereoscopic video data collection?

Stereoscopic video data collection is capturing the same scene through two calibrated cameras spaced like human eyes, so every frame pair carries recoverable depth. For AI it produces training data for depth estimation, 3D hand and object pose, and manipulation policies that need to judge distance the way a person does.

How it works

What makes stereo data trainable.

01

Per-session calibration

Intrinsics, extrinsics, and the stereo baseline are measured every session and shipped with the episode, so rectification is exact rather than assumed.

02

Hardware-synchronized pairs

Left and right frames trigger on one clock alongside depth, IMU, and audio, so there is no drift to correct between eyes or between sensors.

03

Human-baseline geometry

A lens spacing close to human interpupillary distance matches the viewpoint humanoid and head-mounted robot cameras actually see from.

04

Labels in 3D, not just 2D

3D hand joints, object poses, and contact events are annotated against the stereo geometry, not guessed from a single frame.

At a glance

What you get.

Streams
Stereo RGB · depth · IMU · audio
Calibration
Per session, shipped with episode
Annotation
3D hand pose · object pose · contacts
Spec to first batch
a median of 48 hours (deep coverage)
License
Perpetual, buyer-owned

Every collection runs through the same seven-stage end-to-end custom collection pipeline. End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline.

FAQ

Stereoscopic video data collection, answered.

Why stereo instead of a single camera plus depth sensor?
Stereo gives dense depth that is aligned pixel-for-pixel with the RGB view and works outdoors where active depth sensors struggle. Many programs capture both, and we sync them on one clock.
Do you deliver rectified or raw stereo?
Both, if you want. Raw frames ship with calibration so you can re-rectify, and rectified pairs plus disparity can be added to the delivery.
Can stereo capture be combined with wrist cameras?
Yes. Head-mounted stereo, wrist views, and a third-person reference camera can all trigger on the same clock in one episode.

Brief the gap. Get data back fast.

Tell us what your model is missing. We will quote reach, timeline, and price against a written spec, with a first validated batch in a median of 48 hours where coverage is deep.