How it works

From raw prompts to a dataset you can defend.

Six stages. Each one produces a record, which is why the finished dataset can be traced back to the person who made each judgement and the standard they were held to.

The pipeline

What happens to a batch.

Ordered, because it genuinely is: nothing is reviewed before it is rated, and nothing ships before it is reviewed.

01

Upload and configure

A batch arrives as CSV or Excel — prompts with candidate responses. You set the rubric, the pay rate, the time ceiling, how much of the batch is overlapped and how many raters see each overlapped row.
02

Qualify the cohort

Raters sit the project's exam, scored against known answers. Only those above the pass mark receive live work, and batches can be restricted further to higher rater levels.
03

Rate

Tasks are handed out one at a time and locked to a single rater, so two people can never be given the same row. Gold-standard tasks are seeded in invisibly.
04

Measure

Gold answers score accuracy continuously. Overlapped rows are compared for agreement, and groups where raters diverge are flagged as disputed rather than averaged into a false consensus.
05

Review

A reviewer sees completed work in a queue with the rating, the justifications and both responses, and returns one of three verdicts.
06

Deliver

Accepted work exports as JSONL preference pairs or a full CSV, alongside the quality figures for the batch.

Review outcomes

Three verdicts, three different consequences.

The distinction matters. Treating a weak justification the same as a wrong answer either punishes people unfairly or lets bad data through. Pick one to see what it does.

Rating
Stands as submitted
Payment
Kept
Task
Enters the delivered dataset
Rater
Notified that it passed

Rater’s ledger · example ₹50 task

Task paymentTASK_PAYMENT+₹50
Net for this task₹50

Measurement

How agreement is scored.

Overlap is only worth paying for if something is computed from it.

Majority verdict

The most common call across the raters who saw the row, with the share that chose it.

Pairwise agreement

The proportion of rater pairs that landed on the same call — the plain-language figure.

Chance-corrected agreement

Fleiss' kappa across the batch, so agreement that would happen by luck alone is not counted as signal.

Distance on disagreement

How far apart raters were on the preference scale, which separates a near-miss from an opposite call.

Try it — set where each rater landed.

Overlap group · 3 ratersDisputed
Rater 1
Rater 2
Rater 3
Majority
A preferred
2 of 3 raters
Pairwise agreement
0.33
Mean distance
2.7 steps
Genuinely different calls

Verdicts are grouped as A, tie or B before comparing, so “A is better” and “A is much better” count as agreement — the distance figure is what tells them apart. Groups below 0.50 pairwise agreement are flagged for a human.

Delivery

What lands in your hands.

Two formats, both drawn from the same reviewed records.

Preference pairs — JSONL. One record per accepted task: the prompt, the chosen and rejected response, a graded reward score and the rater's written justification. Ties are left out — a tie expresses no preference, and training on it teaches one that was never given.

Full record — CSV. Every task in the batch with its rating, rubric answers, review verdict, reviewer comment and time taken, so you can see what made it through and what didn't.

batch_7f3a9c_rlhf.jsonl
{
"task_id": "9F2C-014",
"prompt": "Explain why the sky is blue — for a ten-year-old.",
"chosen_response": "Sunlight looks white, but it’s really every colour…",
"rejected_response": "Rayleigh scattering. Scattered intensity is…",
"reward_score": 2,
"human_justification": "A pitches it at the reader’s level; B is correct but not for a child."
}
  • 3 / 2 / 1 — much / better / slightly
  • Ties are left out: no preference was expressed.
  • Work returned or failed in review is excluded.

Want to see it against your own data?

Send a sample and we will run it through as a pilot batch, quality figures included.