LiveQuality measured on every task

Human judgement, measured at scale.

Cicada runs rubric-driven evaluation for AI teams. Every label is scored against gold answers, checked for agreement between raters, and delivered with the trail that proves it.

Try it — judge a real task the way our raters do.

Task 9F2C-014Gold task · hidden from rater

Prompt

Explain why the sky is blue — for a ten-year-old.

AResponse A

Sunlight looks white, but it’s really every colour mixed together. When it hits the air, tiny gas molecules bounce blue light around far more than red — so blue reaches your eyes from every part of the sky.

BResponse B

Rayleigh scattering. Scattered intensity is proportional to 1/λ⁴, so shorter wavelengths dominate diffuse skylight and the sky appears blue.

Which response is better?

Press 1–7 or click

Make your call. Three other raters judged this task independently — you’ll see where they landed, and how the platform scores the group.
Pairwise preferenceCustom rubricsGold-standard checksOverlap samplingFleiss’ κ agreementQA review & reworkAppealsQualification examsRater tiersAppend-only ledgerFull audit logJSONL export

The loop

Six steps between a prompt and a label you can trust.

Every control exists because unmeasured labelling fails quietly — and you find out from the model, months later.

  1. 01

    Define the rubric

    Criteria, scales and required justifications are set per batch and validated on the way in — so an incomplete rating can’t enter the dataset.

    Rubric · batch ALPHA-WORKValidated
    Overall preference7-point scalerequired
    Factual accuracy1–5 scalerequired
    Task completion1–5 scalerequired
    JustificationFree textrequired

    Every submission is checked against this schema server-side.

  2. 02

    Qualify the cohort

    Raters pass an exam built from gold answers before touching live work. Harder batches open only to higher tiers, and tiers move with measured accuracy.

    Qualification exam · Alpha ProjectPassed

    Gold answers matched

    8/10

    Pass mark 70%

    L1
    L2
    L3
    L4
    L5

    Passed — live work on this project is now in the queue.

  3. 03

    Rate blind and independent

    Each rater sees the prompt, both responses and the rubric — never anyone else’s answer. Every required field and a written justification must be in before a rating can be submitted.

    Task 9F2C-014 · RatingBlind

    “Explain why the sky is blue — for a ten-year-old.”

    A
    B
    A
    A
    A
    =
    B
    B
    B
    A pitches it at the reader’s level; B is correct but unreadable for a child
  4. 04

    Measure while it runs

    Known-answer tasks are mixed invisibly into the queue, and a share of every batch is rated by several people. Accuracy and agreement are live numbers, not a post-mortem.

    Quality · liveHealthy

    Gold accuracy

    93.4%

    Fleiss’ κ

    0.71

    2 overlap groups disputed this batch → routed to adjudication.

  5. 05

    Review before it counts

    A reviewer passes, returns or fails each rating. Failed work is withdrawn and its payment reversed on the ledger — and the rater can appeal with a reason.

    Review queue3 decided

    Passed 7A1C-031

    Accurate and well justified.

    Needs rework 88F0-112

    Expand the accuracy justification with a source.

    Failed 2D9B-078

    Justification contradicts the choice. −₹50 reversed.

  6. 06

    Deliver with the trail

    Export preference pairs as JSONL for training, or the full record as CSV — every rating, justification, rater and QA verdict behind each label.

    export · ALPHA-WORK.jsonlReady
    { "task_id": "9F2C-014",
    "prompt": "Explain why the sky is blue…",
    "chosen_response": "Sunlight looks white, but…",
    "rejected_response": "Rayleigh scattering…",
    "reward_score": 2,
    "human_justification": "A is pitched at the reader…" }
    JSONL pairsFull CSV

Why it holds up

Most annotation is trusted.
Ours is measured.

Accuracy, agreement, review and money all leave a record — so the data can be questioned, and answer back.

Gold-standard scoring

Known-answer tasks are mixed invisibly into ordinary work, so accuracy is sampled on every batch — not discovered at review time.

Accuracy on gold tasks · example project

0.0%

+31.4 pts over 12 batches

Measured agreement

Overlapping work gets a chance-corrected agreement score. Disputes go to a human instead of being averaged away.

κ = 0.82Agreed
R1
R2
R3
A++=B++

An audit trail for everything

Role changes, reviewer edits, gold answers and money movements — recorded, with what changed.

nowQA_EDITED_RATINGqa_alpha · 7A1C-031 · before/after kept
2s agoPAYOUT_PROCESSED₹1,200 · t_rater1
4s agoAPPEAL_RESOLVEDupheld · payment restored
6s agoROLE_CHANGEDt_rater4 · rater → qa

A ledger, not a counter

Every credit, reversal and payout is an append-only entry. Nothing gets paid twice; nothing silently vanishes.

Task payment9F2C-014+₹50
QA reversal2D9B-078−₹50
Appeal restored2D9B-078+₹50
PayoutUPI ••66−₹100

Access that’s earned

Raters qualify by exam and move between five tiers on measured quality. Hard batches open to proven people.

L1
L2
L3
L4
L5
0Point preference scaleStrength of preference, not just a winner.
0×Raters per overlap itemAgreement measured, not assumed.
0Rater tiersAccess earned through measured accuracy.
0%Decisions loggedEvery verdict, edit and payment.

Two sides, one standard

A marketplace only works if both sides are held to something.

Evaluation data you can defend in a review.

Bring a rubric or let us build one. We assemble a qualified cohort, sample for agreement, review the output and hand back labels with the evidence behind them.

  • Custom rubrics per batch, validated on every submission
  • Gold tasks scoring accuracy continuously
  • Agreement measured on overlapping work
  • JSONL preference pairs or full CSV exports

Batch ALPHA-WORK

Example delivery summary

Ready to ship
Labels delivered
2,400
Gold accuracy
93.4%
Agreement (κ)
0.71
Reviewed by QA
100%
Disputes adjudicated
12

Make your evaluation data defensible.

Send a sample of your data and the judgement you need made. We’ll come back with a rubric, a cohort plan and what quality assurance will look like.