Solutions / Text & language
Text written by the people you are modeling.
Prompts, conversations, translations, and domain-specific writing authored by native speakers — for training, fine-tuning, and evaluating language models in the languages and registers your users actually write in.

Definition
Text data collection
Text data collection is the authoring or gathering of written language from real contributors against a defined spec — prompts, dialogue, translations, or domain text — with the rights to use it recorded up front.
What we collect
The formats teams ask us for.
Not an exhaustive menu — if what you need is not here, it is a custom collection, which we also do.
Prompts & instructionsHuman-written prompts and instructions across tasks and difficulty levels for instruction tuning.
Conversations & dialogueMulti-turn dialogue, including role-played support and assistant conversations.
Translation & localizationHuman translation and localization by native speakers, not machine output cleaned up.
Domain & expert textWriting from contributors with real domain knowledge — legal, medical, technical, financial.
Preference & ranking dataHuman preference judgements and rankings for alignment and RLHF-style training.
Red-team & edge promptsAdversarial and edge-case prompts authored to probe model behaviour.
What shapes the spec
The decisions we settle before collecting.
- Languages & locales
- 500+ languages and locales
- Author sourcing
- Native speakers, screened for domain expertise where required
- Task design
- Guidelines, difficulty tiers, and inter-annotator agreement targets set in the spec
- Delivery
- Structured text (JSONL/CSV) with author metadata, language tags, and QA labels
- Consent basis
- Contributor agreement granting commercial usage rights
Where it is used
What teams train with it.
- LLM instruction tuning and fine-tuning
- Human preference and alignment (RLHF-style) data
- Machine-translation training and evaluation
- Domain adaptation with expert-authored text
- Red-teaming and safety evaluation sets
How we run it
The standards behind every batch.
These apply to every modality — they are the reason the data is usable rather than merely large.
- Collected to a written spec
- Nothing is gathered speculatively. We agree the target — languages, demographics, devices, environments, edge cases — before a single contributor is briefed.
- Contributors paid hourly
- People are paid for their time, including setup and retakes, not a bounty per item. A per-item rate optimises for volume and quietly wrecks quality.
- Documented, informed consent
- Every contributor signs a release granting the usage rights you need before collection begins. You receive the consent artefacts with the batch.
- Discard rather than downgrade
- If an item cannot meet the spec or a bystander cannot be de-identified, it is dropped — not shipped at a discount to pad the count.
- Multi-layer QA with a visible reject log
- Automated checks plus human review, and you see the reject reasons, not just the accepted items.
- Buyer-owned commercial license
- You receive a perpetual, buyer-owned license with a data card recording jurisdictions of capture and the consent basis.
FAQ
Text & language collection, answered.
- Is the text human-authored or model-generated?
- Human-authored by default. When you need human prompts, translations, preference judgements, or domain writing, that is what we collect — from screened native speakers and domain experts, not machine output lightly edited. If you specifically want model-in-the-loop data, we design that explicitly rather than blurring the line.
- Which languages can you cover?
- Sourcing runs through a vetted global crowd across 500+ languages and locales, so we can author and translate well beyond the usual high-resource set. Under-served languages take longer to staff, and we give you the realistic timeline before you commit.
- Can you collect expert or domain-specific text?
- Yes. We screen contributors for real domain knowledge — legal, medical, technical, financial — and set inter-annotator agreement and review targets in the spec so the output holds up to expert scrutiny.
Scope a text & language collection.
Bring your spec or your problem. You will get a scoped estimate — reach, timeline, and price — before any commitment.