Guide

Why Failure Cases Matter in Robot Learning

Why deliberately capturing transit, occlusion, low light and recovery attempts is what makes training data teach robustness instead of a best-case demo.
By the Firsthand capture teamLast updated September 26, 2026

Short answer

Failure cases matter because policies fail exactly where highlight-reel training data cuts: occlusion, low light, drops, and recovery. Failure-case coverage is the deliberate inclusion of these hard parts alongside clean successes, so a model actually learns to recover instead of breaking the first time reality is messy. Firsthand ships these moments uncut.

Last reviewed:

What counts as a failure case?

A failure case is any part of a task that is not a clean, uninterrupted success: transit between locations, occlusion during a grasp, low light, a dropped object, or a recovery attempt after something goes wrong. It is distinct from a failed episode — the episode itself can still be validated and delivered, it just contains the messy middle instead of only the clean outcome.

Failure-case coverage is the deliberate inclusion of these hard parts in a dataset, rather than editing them out in favor of only best-case demonstrations.

Why do policies need failure cases to learn from?

Policies fail in exactly the situations that highlight reels cut. If training data only shows clean grasps in good light with nothing in the way, the model never sees what recovery looks like, and it breaks the first time reality is messy — which, outside a demo, is most of the time.

This is the difference between a dataset that teaches robustness and one that teaches a best-case demo the policy can only reproduce under the same clean conditions it was trained on.

How does failure-case coverage get captured and delivered?

Firsthand ships the transitions and hard cases uncut — walking between rooms, exposure hunting in dim light, occluded grasps — because that is what a head-mounted camera actually returns when it is not curated down to highlights.

Clean-success-only clips
Fast to review, but the model never sees recovery or messy conditions
Failure-case coverage
Transit, occlusion, low light and recovery included; teaches robustness, not just a best case
Validated hour
Either kind of episode can be validated — coverage is about what is included, not whether it passed QA

Curated vs. failure-case-covered training data

How do you spec failure-case coverage into a custom collection?

A skill spec is where failure-case coverage gets planned rather than left to chance: naming the occlusion angles, lighting conditions, and recovery scenarios that matter for the target task before capture starts, so those cells count toward validated hours instead of being an afterthought.

FAQ

Why Failure Cases Matter in Robot Learning, answered.

01

Does a failure case mean the episode failed quality validation?

No. A failure case describes what happens within the task — occlusion, low light, a drop, a recovery — not whether the episode passed validation. A messy episode can still be a validated hour if it meets sync, calibration and consent requirements.

02

Why not just train on clean successful demonstrations?

Because deployment is not clean. If a model only ever sees best-case grasps, it has no signal for what to do when something goes wrong, and it fails the first time reality departs from the demo.

03

How does Firsthand avoid editing out the hard parts?

Firsthand ships the transitions and hard cases uncut — walking between rooms, exposure hunting in dim light — rather than curating footage down to only the clean, successful segments.

04

How do I make sure failure cases are covered in a custom dataset?

Name the specific occlusion angles, lighting conditions and recovery scenarios that matter for your task in the skill spec before capture starts, so they are planned coverage cells rather than left to chance.

05

Does failure-case coverage slow down delivery?

It depends on scope: shallow, well-covered cells can still ship in a median of 48 hours, while long-tail failure-case coverage that requires more scenario variety takes longer. See how validated hours are counted for the full breakdown.

Check it against the sample pack.

40 episodes across 4 environments, delivered in the exact schema these guides describe.