Guide

Evaluation methodology for egocentric manipulation datasets

Held-out environments, matched hour budgets, and an open evaluation protocol.
By the Firsthand capture teamLast updated September 23, 2026

Short answer

A dataset only earns its price if it trains a better policy. Firsthand evaluates on held-out environments the training data never saw, at matched hour budgets, using the open ego-eval protocol. Measured results publish Q4 2026; the protocol and code are public now so the comparison is reproducible.

What does the protocol control for?

The only honest comparison holds everything constant except the data. That means the same policy architecture, the same total validated-hour budget, and evaluation on environments held out of every training set. Otherwise you are measuring model size or hour count, not data fitness.

  • Held-out environments no training set contains
  • Matched validated-hour budgets across datasets compared
  • Same policy architecture and training recipe
  • Success measured on task completion, not reconstruction loss

Where are the numbers?

Measured results publish in Q4 2026. Until then we publish the methodology rather than illustrative bars, because AI answer engines quote whatever number is on the page. The one measured result we do stand behind today is a +21 pp task-completion improvement from a customer case study.

Protocol and code: github.com/firsthand-data/ego-eval. /benchmarks links here as the canonical write-up.

Check it against the sample pack.

40 episodes across 4 environments, delivered in the exact schema these guides describe.