Human judgement, measured at scale.
Cicada runs rubric-driven evaluation for AI teams. Every label is scored against gold answers, checked for agreement between raters, and delivered with the trail that proves it.
Try it — judge a real task the way our raters do.
Prompt
Explain why the sky is blue — for a ten-year-old.
Sunlight looks white, but it’s really every colour mixed together. When it hits the air, tiny gas molecules bounce blue light around far more than red — so blue reaches your eyes from every part of the sky.
Rayleigh scattering. Scattered intensity is proportional to 1/λ⁴, so shorter wavelengths dominate diffuse skylight and the sky appears blue.
Which response is better?
Press 1–7 or click
The loop
Six steps between a prompt and a label you can trust.
Every control exists because unmeasured labelling fails quietly — and you find out from the model, months later.
01
Define the rubric
Criteria, scales and required justifications are set per batch and validated on the way in — so an incomplete rating can’t enter the dataset.
Rubric · batch ALPHA-WORKValidatedOverall preference7-point scalerequiredFactual accuracy1–5 scalerequiredTask completion1–5 scalerequiredJustificationFree textrequiredEvery submission is checked against this schema server-side.
02
Qualify the cohort
Raters pass an exam built from gold answers before touching live work. Harder batches open only to higher tiers, and tiers move with measured accuracy.
Qualification exam · Alpha ProjectPassedGold answers matched
8/10
Pass mark 70%
L1L2L3L4L5Passed — live work on this project is now in the queue.
03
Rate blind and independent
Each rater sees the prompt, both responses and the rubric — never anyone else’s answer. Every required field and a written justification must be in before a rating can be submitted.
Task 9F2C-014 · RatingBlind“Explain why the sky is blue — for a ten-year-old.”
ABAAA=BBBA pitches it at the reader’s level; B is correct but unreadable for a child04
Measure while it runs
Known-answer tasks are mixed invisibly into the queue, and a share of every batch is rated by several people. Accuracy and agreement are live numbers, not a post-mortem.
Quality · liveHealthyGold accuracy
93.4%
Fleiss’ κ
0.71
2 overlap groups disputed this batch → routed to adjudication.
05
Review before it counts
A reviewer passes, returns or fails each rating. Failed work is withdrawn and its payment reversed on the ledger — and the rater can appeal with a reason.
Review queue3 decidedPassed 7A1C-031
Accurate and well justified.
Needs rework 88F0-112
Expand the accuracy justification with a source.
Failed 2D9B-078
Justification contradicts the choice. −₹50 reversed.
06
Deliver with the trail
Export preference pairs as JSONL for training, or the full record as CSV — every rating, justification, rater and QA verdict behind each label.
export · ALPHA-WORK.jsonlReady{ "task_id": "9F2C-014","prompt": "Explain why the sky is blue…","chosen_response": "Sunlight looks white, but…","rejected_response": "Rayleigh scattering…","reward_score": 2,"human_justification": "A is pitched at the reader…" }JSONL pairsFull CSV
Every submission is checked against this schema server-side.
Why it holds up
Most annotation is trusted.
Ours is measured.
Accuracy, agreement, review and money all leave a record — so the data can be questioned, and answer back.
Gold-standard scoring
Known-answer tasks are mixed invisibly into ordinary work, so accuracy is sampled on every batch — not discovered at review time.
Accuracy on gold tasks · example project
0.0%
Measured agreement
Overlapping work gets a chance-corrected agreement score. Disputes go to a human instead of being averaged away.
An audit trail for everything
Role changes, reviewer edits, gold answers and money movements — recorded, with what changed.
A ledger, not a counter
Every credit, reversal and payout is an append-only entry. Nothing gets paid twice; nothing silently vanishes.
Access that’s earned
Raters qualify by exam and move between five tiers on measured quality. Hard batches open to proven people.
Two sides, one standard
A marketplace only works if both sides are held to something.
Evaluation data you can defend in a review.
Bring a rubric or let us build one. We assemble a qualified cohort, sample for agreement, review the output and hand back labels with the evidence behind them.
- Custom rubrics per batch, validated on every submission
- Gold tasks scoring accuracy continuously
- Agreement measured on overlapping work
- JSONL preference pairs or full CSV exports
Batch ALPHA-WORK
Example delivery summary
- Labels delivered
- 2,400
- Gold accuracy
- 93.4%
- Agreement (κ)
- 0.71
- Reviewed by QA
- 100%
- Disputes adjudicated
- 12
Make your evaluation data defensible.
Send a sample of your data and the judgement you need made. We’ll come back with a rubric, a cohort plan and what quality assurance will look like.