About
Models are judged by people. Judge the judging.
Human preference sits underneath most of how modern models are aligned and assessed. It is usually collected with far less rigour than the training run that consumes it.
Why we built this
Unmeasured labelling fails quietly.
A rater having a bad week, a rubric two people read differently, a batch where everyone silently converged on the wrong convention — none of it announces itself. It arrives later, as a model that behaves oddly in a way nobody can trace.
The usual answer is to review a sample after the fact. That finds the worst of it, late, and tells you nothing about the work you did not sample.
So we built the measurement into the work itself: known-answer tasks mixed invisibly into the queue, deliberate overlap between raters, and a review step that has real consequences. The result is a dataset that arrives with its own evidence.
Principles
What we hold to.
Measure, don't assert
Any vendor can claim high quality. We would rather show the accuracy scores, the agreement figures and the rejections. If a number is not computed from the work itself, it is marketing.
Both sides are the customer
A rating platform that treats its raters badly produces bad ratings — carelessly, and quite fast. Clear rubrics, explained rejections and an honest ledger are quality controls, not perks.
Disagreement is information
When three people read the same pair and land somewhere different, that is worth knowing. Averaging it away produces a confident number and destroys the finding underneath it.
Everything leaves a record
Role changes, reviewer edits, deletions, payments and appeals are all written down, including what the data looked like beforehand. Work you cannot reconstruct is work you cannot defend.
Disagreement is information
The average is where the finding goes to die.
Three raters, one pair of responses. Flip between averaging their calls and keeping them — one version looks more confident; the other is telling you something.
Disputed — flagged for a human
Two raters preferred A; one clearly preferred B. That split is the finding, so it goes to a reviewer instead of into the dataset.
The platform
Built around the controls, not bolted on.
Work with us, or work for us.
We are looking for AI teams with evaluation to run, and for careful people to do it.