For clients

Evaluation data with the evidence attached.

You should not have to take a vendor's word for label quality. Every batch we return carries the accuracy scores, agreement figures and review history behind it.

Capabilities

Tell us the judgement you need made.

If it can be written as a rubric, it can be run as a batch. Most work falls into three shapes.

Pairwise preference

Two candidate responses to the same prompt, ranked on a seven-point scale with a written justification. The standard shape for reward modelling.

Rubric scoring

Your own criteria, defined per batch — scales and free text. Validated on submission so an incomplete rating never enters the set.

Correctness review

Factual accuracy, instruction-following and safety judgements against guidelines you supply, with the reasoning captured alongside the verdict.

Build a rubric in the open

Toggle criteria to see the schema we validate and the form your raters get.

Batch rubricValid · 3 required fields

Criteria

[
{ "id": "helpfulness", "type": "slider", "min": 1, "max": 5 },
{ "id": "factual_accuracy", "type": "slider", "min": 1, "max": 5 },
{ "id": "rationale", "type": "text" }
]

What the rater sees · per response

Helpfulness

12345

Factual accuracy

12345

Rationale

Required before submitting…

Quality controls

Four checks, running the whole time.

None of these are periodic audits. They run on the work as it happens, which is the only way a problem surfaces while it is still cheap to fix.

Gold accuracy93%

Measured continuously on known-answer tasks seeded into ordinary work, so a drift in quality shows up while it is still cheap to correct.

01

Qualification

Raters sit a scored exam on your project's rubric before they touch live work. Below the pass mark, they do not get any.

02

Gold-standard sampling

Tasks with known answers are mixed invisibly into ordinary work. Accuracy is therefore measured continuously, on work nobody knew was being marked.

03

Overlap and agreement

A configurable share of each batch is rated independently by several people. Where they diverge, the group is flagged rather than averaged.

04

Second-pass review

A reviewer sees completed work before it counts. They can accept it, return it to its author with a reason, or reject it — which withdraws the payment.

What you receive

Data in the shape you will train on.

JSONL preference pairs

Chosen and rejected responses with a graded reward score, ready for reward-model training. Ties are excluded — they carry no preference.

Full CSV export

Every task, rating, justification, rubric answer and review verdict for the batch.

Quality reporting

Per-rater accuracy against gold tasks, agreement scores on overlapping work, and throughput by project and batch.

batch_7f3a9c_rlhf.jsonl
{
"task_id": "9F2C-014",
"prompt": "Explain why the sky is blue — for a ten-year-old.",
"chosen_response": "Sunlight looks white, but it’s really every colour…",
"rejected_response": "Rayleigh scattering. Scattered intensity is…",
"reward_score": 2,
"human_justification": "A pitches it at the reader’s level; B is correct but not for a child."
}
  • 3 / 2 / 1 — much / better / slightly
  • Ties are left out: no preference was expressed.
  • Work returned or failed in review is excluded.

Access and confidentiality

Work is scoped to the project it belongs to. Raters and reviewers see only their own project's data, payout details are encrypted with authenticated encryption, and every privileged action — role changes, deletions, reviewer edits, reads of bank details — is written to an immutable log.

Sizing and turnaround

Batches are uploaded as CSV or Excel and fan out to a qualified cohort immediately. Throughput scales with the size of the cohort you fund; overlap and review depth are configurable per project, so you can trade cost against confidence deliberately.

Send us a sample.

A hundred rows and a description of the judgement you need is enough. We will come back with a draft rubric, a cohort plan and what the quality controls would look like for your data.

Talk to us