Custom collection
End-to-end custom data collection for AI.
- Reach
- 150+ countries
- Languages
- 500+ locales
- Modalities
- Any, or combined
- License
- Buyer-owned

What it is
One pipeline, problem to owned dataset.
End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline — discovery, spec, sourcing, capture, annotation, QA, and licensed delivery — so you brief one team instead of stitching together crowds, tools, and reviewers yourself.
An off-the-shelf dataset answers a question someone else asked. Custom collection answers yours — and end-to-end means you do not have to become a general contractor to get it. Instead of hiring a crowd here, a labelling shop there, and a QA tool to referee, you agree one spec with one team that carries it all the way to a licensed delivery.
That is the whole difference: the people who write the spec are the people who ship the data, so the accept criteria and the consent chain are enforced at every stage rather than lost between vendors.
Capture
From the phone in a pocket to a full rig.
First-person capture does not have to start with specialist hardware. A contributor can shoot to spec on an iPhone in a simple chest mount, and scale up to the multi-camera rig when the project needs wrist views, depth and pose. Same spec, same QA bar, either way.

Starting on commodity phones is how a program reaches people fast and anywhere in the world. Anyone we recruit already owns the camera; the spec tells them how to mount it, frame the task, and what counts as a usable take.
Want to capture in your own town? Become a collector →
Real first-person clips
The same footage example from our homepage — unedited first-person captures, switch between them.
- Frame
- 960 × 540
- Length
- 8 s
- File
- 4.7 MB
- Audio
- Present in source, muted here
A transit segment — carrying a tool between floors. Exactly the footage staged datasets cut, and a reason policies fail to move between rooms. The image circle is letterboxed into a 16:9 canvas, so the black margins are the proxy container, not lost data.
The pipeline
Seven stages, one owner.
Every custom collection runs through the same accountable pipeline. Each stage has an owner, an output, and a bar it has to clear before the next one starts.
- 01DiscoveryWe start from the failure, not a schema. You describe what you are training and where it breaks.
- 02SpecificationThe problem becomes a written, testable spec you approve before anyone is briefed.
- 03Sourcing & staffingWe recruit and screen the exact contributors the spec calls for, wherever they are.
- 04CaptureContributors are briefed, paid hourly, and collected against the spec — with the right rig where it matters.
- 05Annotation & structuringRaw capture is labelled, transcribed, or structured into the schema your loader expects.
- 06QA & anonymizationAutomated checks plus human review against the accept criteria, with a visible reject log.
- 07Licensed deliveryData ships in your format with the consent artefacts and a buyer-owned commercial license.
Each stage broken down — what happens, what you get, and how it is measured — on the end-to-end pipeline page.
For AI training
Built for the model you are training.
Custom collection is not one service — it is the same pipeline pointed at whatever your system needs. The common cases:
The full breakdown by model type, with what we collect and an example brief for each, is on custom data collection for AI training.
Where custom collection starts
A custom spec is a conversation about the thin cells.
This is the same coverage picture we publish on the homepage, gaps included. The deep bars are shipping now; the thin, dashed ones are exactly where a custom collection is scoped — and where a pilot buys you exclusivity on the capture.
Dataset diversity
Generalization is a spread you scope on purpose.
Models fail on the world they never saw. A custom collection balances three axes at once — the activity, the environment it happens in, and where in the world it was captured — so the set does not quietly collapse onto whoever was easiest to recruit.
Volume is the easy number to sell. Diversity is the one that decides whether a policy survives contact with a real kitchen it has never seen. Every Firsthand program is weighted to your spec, so these proportions move — what stays fixed is that the spread across activity, scene, and geography is a deliberate decision written into the spec, not an accident of who signed up first.
Switch the axes to see how one representative program balances out. Thin a segment on purpose for a focused model, or widen it for one that has to generalize — either way you approve the mix before capture begins.
The per-cell version of this — every environment against every condition — is the coverage matrix above. Diversity here, depth there.
Distribution tells you the shape of the set; the scatter tells you it is not one person capturing the same corner a thousand times. Each dot is a contributor, placed by how many distinct scenes they covered against the validated hours they delivered, colored by region.
A healthy program is a wide, up-and-to-the-right cloud — broad scene coverage from every region, not a tight cluster. Thin spots are where we open recruitment before the set locks.
Program velocity
Volume ramps; the reject gap closes.
A custom collection is not a one-shot dump. Capture builds week over week, and the distance between what is submitted and what clears QA shrinks as contributors internalise the spec.
The dashed line is everything contributors turn in; the filled line is what passes review. Early on the gap is wide — that is the spec doing its job, rejecting takes that miss a lighting, framing, or consent requirement.
By the back half of the program the two lines converge: the same people now shoot to spec the first time. Rejected hours are recaptured, never quietly shipped to hit a number.
This is the acceptance side of the sync and coverage bars elsewhere on the site — the same QA bar, plotted over time.
Why end-to-end
The failures live in the handoffs.
Most bad AI data is not badly collected — it is well collected and then wrecked in a handoff. Owning the whole chain is how the standard survives.
- One owner for the whole chain
- Sourcing, capture, annotation, QA, and licensing are one accountable pipeline. When something is wrong you have one team to call, not a crowd platform blaming a labeling vendor blaming a tool.
- The spec is enforced at every stage
- Because the same team writes the spec and ships the data, the accept/reject criteria are applied end to end — not lost in a handoff between the people who collected it and the people who labelled it.
- Provenance survives the whole journey
- Consent is captured at sourcing and travels with the item through delivery. A stitched-together pipeline is where consent chains break; an end-to-end one is where they hold.
- You get data, not a project to manage
- No orchestrating vendors, reconciling formats, or chasing who owns the license. You brief once and receive an owned, structured, documented dataset.
FAQ
Custom collection for AI, answered.
- What is end-to-end custom data collection?
- End-to-end custom data collection is a managed service that takes an AI data need from problem to owned dataset in a single accountable pipeline — discovery, spec, sourcing, capture, annotation, QA, and licensed delivery — so you brief one team instead of stitching together crowds, tools, and reviewers yourself.
- What does "custom collection for AI" actually deliver?
- A dataset built to your specification rather than pulled off a shelf — the exact modality, languages, demographics, devices, environments, and edge cases your model needs, collected by paid contributors with signed consent, structured into your delivery format, and handed over under a perpetual, buyer-owned commercial license with a visible reject log and a data card.
- Why buy the whole pipeline instead of assembling it yourself?
- Because the failures in AI data live in the handoffs. A crowd platform, a labelling vendor, and a QA tool each optimise their own step; the spec and the consent chain get lost between them. End-to-end means one team owns sourcing, capture, annotation, QA, and licensing, so the accept criteria are enforced the whole way through and there is one owner accountable for the result.
- Do I need a finished spec before I start?
- No. Most projects begin from a gap — a model that fails in a certain language, region, or condition — not a schema. Turning that problem into a written spec with explicit accept and reject criteria is the first stage of the work, and you approve it before any collection begins.
- Which types of AI can you collect data for?
- LLMs and assistants, speech and ASR, computer vision, embodied AI and robotics, multimodal and sensor models, and safety or evaluation sets. Egocentric first-person video is our core specialism; the other modalities run through the same spec-first, consent-documented pipeline.
Describe the data you wish existed.
You do not need a finished spec — just the problem. You will get a scoped estimate with reach, timeline, and price before any commitment, and one team accountable from spec to delivery.