For AI teams

Human data from people you can verify.

Held-out evaluation sets, RLHF and preference data, expert human evaluation and ground truth, produced only by vetted experts, with provenance for every item.

Evaluations

Held-out benchmarks

New, unpublished test sets that expose where your model fails, with every error labelled.

Post-training

RLHF & preference data

Rubric-based ratings, pairwise preferences and expert rewrites in your target languages and fields.

Ground truth

Transcription & annotation

OCR/HTR ground truth, translation and domain annotation, typed blind and adjudicated.

Evidence first

On our public Urdu Nastaliq sample, frontier models made a mistake in about 70% of lines. Read the method and results.

Request a proposalHow vetting works

Questions

Can we start with a pilot?

Yes. Most teams start with a fixed-price pilot, then scale once quality is proven.

How do you prove data quality?

Every item carries provenance: which vetted expert produced it, when, under which guideline version, and how it was reviewed. We share our method and sample results openly.

Which languages and fields?

Verified experts across 135 languages and 28 fields, with particular depth in South Asian languages and scripts.