How it works
From raw prompts to a dataset you can defend.
Six stages. Each one produces a record, which is why the finished dataset can be traced back to the person who made each judgement and the standard they were held to.
The pipeline
What happens to a batch.
Ordered, because it genuinely is: nothing is reviewed before it is rated, and nothing ships before it is reviewed.
Upload and configure
Qualify the cohort
Rate
Measure
Review
Deliver
Review outcomes
Three verdicts, three different consequences.
The distinction matters. Treating a weak justification the same as a wrong answer either punishes people unfairly or lets bad data through. Pick one to see what it does.
- Rating
- Stands as submitted
- Payment
- Kept
- Task
- Enters the delivered dataset
- Rater
- Notified that it passed
Rater’s ledger · example ₹50 task
Measurement
How agreement is scored.
Overlap is only worth paying for if something is computed from it.
Majority verdict
The most common call across the raters who saw the row, with the share that chose it.
Pairwise agreement
The proportion of rater pairs that landed on the same call — the plain-language figure.
Chance-corrected agreement
Fleiss' kappa across the batch, so agreement that would happen by luck alone is not counted as signal.
Distance on disagreement
How far apart raters were on the preference scale, which separates a near-miss from an opposite call.
Try it — set where each rater landed.
- Majority
- A preferred
- 2 of 3 raters
- Pairwise agreement
- 0.33
- Mean distance
- 2.7 steps
- Genuinely different calls
Verdicts are grouped as A, tie or B before comparing, so “A is better” and “A is much better” count as agreement — the distance figure is what tells them apart. Groups below 0.50 pairwise agreement are flagged for a human.
Delivery
What lands in your hands.
Two formats, both drawn from the same reviewed records.
Preference pairs — JSONL. One record per accepted task: the prompt, the chosen and rejected response, a graded reward score and the rater's written justification. Ties are left out — a tie expresses no preference, and training on it teaches one that was never given.
Full record — CSV. Every task in the batch with its rating, rubric answers, review verdict, reviewer comment and time taken, so you can see what made it through and what didn't.
{"task_id": "9F2C-014","prompt": "Explain why the sky is blue — for a ten-year-old.","chosen_response": "Sunlight looks white, but it’s really every colour…","rejected_response": "Rayleigh scattering. Scattered intensity is…","reward_score": 2,"human_justification": "A pitches it at the reader’s level; B is correct but not for a child."}
- 3 / 2 / 1 — much / better / slightly
- Ties are left out: no preference was expressed.
- Work returned or failed in review is excluded.
Want to see it against your own data?
Send a sample and we will run it through as a pilot batch, quality figures included.