About

Models are judged by people. Judge the judging.

Human preference sits underneath most of how modern models are aligned and assessed. It is usually collected with far less rigour than the training run that consumes it.

Why we built this

Unmeasured labelling fails quietly.

A rater having a bad week, a rubric two people read differently, a batch where everyone silently converged on the wrong convention — none of it announces itself. It arrives later, as a model that behaves oddly in a way nobody can trace.

The usual answer is to review a sample after the fact. That finds the worst of it, late, and tells you nothing about the work you did not sample.

So we built the measurement into the work itself: known-answer tasks mixed invisibly into the queue, deliberate overlap between raters, and a review step that has real consequences. The result is a dataset that arrives with its own evidence.

Principles

What we hold to.

01

Measure, don't assert

Any vendor can claim high quality. We would rather show the accuracy scores, the agreement figures and the rejections. If a number is not computed from the work itself, it is marketing.

02

Both sides are the customer

A rating platform that treats its raters badly produces bad ratings — carelessly, and quite fast. Clear rubrics, explained rejections and an honest ledger are quality controls, not perks.

03

Disagreement is information

When three people read the same pair and land somewhere different, that is worth knowing. Averaging it away produces a confident number and destroys the finding underneath it.

04

Everything leaves a record

Role changes, reviewer edits, deletions, payments and appeals are all written down, including what the data looked like beforehand. Work you cannot reconstruct is work you cannot defend.

Disagreement is information

The average is where the finding goes to die.

Three raters, one pair of responses. Flip between averaging their calls and keeping them — one version looks more confident; the other is telling you something.

Three raters, one pair
Rater 1
Rater 2
Rater 3
A++=B++

Disputed — flagged for a human

Two raters preferred A; one clearly preferred B. That split is the finding, so it goes to a reviewer instead of into the dataset.

The platform

Built around the controls, not bolted on.

0RolesAdmin, PM, QA, rater, client, owner
0Review verdictsAccept, rework, reject
0Rater tiersEarned on measured accuracy
0Export formatsJSONL pairs and full CSV

Work with us, or work for us.

We are looking for AI teams with evaluation to run, and for careful people to do it.