Solutions / Custom collection

When the data you need does not exist yet.

Some problems have no dataset to buy. The languages are too rare, the device too specific, the task too unusual, or the consent standard too high for anything already on the market. Custom collection is for those: we design the whole program — sourcing, spec, capture, QA, and delivery — around your need, and run it wherever in the world the data actually lives.
Reach
150+ countries
Languages
500+ locales
Modalities
Any, or combined
License
Buyer-owned
A coordinator with a clipboard briefs a contributor wearing a capture device in a workshop, with other workers behind them.
A bespoke program: we brief, staff, and run collection against a spec you approve — wherever in the world the data lives.

What it is

A program, not a product.

Custom data collection is a bespoke program built to gather training data that does not already exist in the form you need — specified by you, staffed and run by us, delivered with documented consent and a buyer-owned license.

An off-the-shelf dataset answers the question someone else asked. A custom collection answers yours. That is the whole distinction: instead of adapting your problem to the data that happens to exist, we build the data around the problem — the exact languages, demographics, devices, environments, and edge cases your model has to survive.

How it works

Six steps from problem to owned dataset.

Most projects start at step one with a gap, not a schema. Turning that into a spec is part of the work.

01 — Intake
You describe the problem: what you are training, where it fails today, and the data you wish existed. No spec required yet — most projects start from a gap, not a schema.
02 — Spec
We turn that into a written specification: modality, languages, demographics, devices, environments, edge cases, volume, and explicit accept/reject criteria. You approve it before anything is collected.
03 — Sourcing
We recruit and screen contributors from the regions, languages, and profiles the spec calls for, through a vetted global crowd. Thin segments get a realistic timeline, not a hopeful one.
04 — Capture
Contributors are briefed, paid hourly for their time, and collected against the spec — with any specialized rig (for example our synchronized egocentric hardware) where the project needs it.
05 — QA
Automated checks plus human review against the accept criteria. Items that miss are rejected and recaptured, and you see the reject log, not just the accepted set.
06 — Delivery
Data ships in your format with per-item metadata, the consent artefacts, and a data card recording jurisdictions and consent basis. You own a perpetual commercial license to the result.

When to use it

The needs custom collection is built for.

A modality mix no single vendor coversSpeech plus video plus sensor data from the same sessions, aligned on one timeline.
A language or region off the beaten pathSpeakers and settings that off-the-shelf datasets simply do not include.
A device or environment constraintSpecific phone models, in-car audio, factory floors, clinical training rooms.
An edge case that decides accuracyThe rare, hard, or adversarial cases your current data is missing.
A demographic balance requirementDeliberate sampling for fairness across age, gender, skin tone, or geography.
A rights and consent standardProvenance clean enough to show your counsel, for regulated or high-stakes use.

If your need maps cleanly to one modality, start from its page — audio, image, video, text, or multimodal.

Worldwide

Run where the data actually lives.

Sourcing runs through a vetted global crowd across 150+ countries and 500+ languages. We state reach qualitatively on purpose: what matters is not a headcount you cannot audit, but whether we can reach the specific speakers, regions, and settings your project needs — and we tell you that, honestly, before you commit.

FAQ

Custom collection, answered.

What is custom data collection?
Custom data collection is a bespoke program built to gather training data that does not already exist in the form you need. Instead of buying a fixed dataset, you specify the modality, languages, demographics, devices, environments, and edge cases, and a collection program is designed, staffed, and run to produce exactly that — with documented consent and a buyer-owned license.
What if I do not have a spec yet?
That is the normal starting point. Most projects begin from a gap — a model that fails in a certain language, region, or condition — not a finished schema. We help turn that problem into a written spec with explicit accept and reject criteria, and you approve it before any collection begins.
Can you collect data anywhere in the world?
Sourcing runs through a vetted global crowd across 150+ countries and 500+ languages, so the practical answer for most projects is yes. Deep segments start quickly; thin or under-served ones — low-resource languages, rare settings, specific demographics — take longer to staff, and we quote a realistic timeline before you commit rather than promising coverage we cannot yet reach.
How do you price a custom collection?
Per engagement, because the cost is driven by things we cannot guess from a page: modality, how thin the coverage is, how dense the annotation must be, and whether you need exclusivity. You get a scoped estimate — reach, timeline, and price — before any commitment, and billing is per accepted item or validated hour. Rejects, recaptures, and setup are not billed.
Who owns the data you collect for me?
You do. Custom collections are delivered under a perpetual, buyer-owned commercial license, with the consent artefacts and a data card recording jurisdictions of capture. Exclusivity is available where you need the data collected for you and no one else.

Describe the data you wish existed.

You do not need a finished spec — just the problem. We will turn it into a scoped estimate with reach, timeline, and price before any commitment.