Designing Rubrics for Expert Model Evaluation
How to turn "that answer is good" into a score two experts would give it the same way - and the biases that break agreement.
July 2026 · 8 min · Wiingy Research
Turning judgment into a number
You have probably sat through two kinds of grading. One teacher handed back essays with marks that felt like a mood ring. An A from her was a B from the guy next door, and nobody could say why. The other handed out a rubric on day one, so you knew an A meant a clear thesis, evidence for every claim, and no hand-waving. Model evaluation is that same fork in the road. Without a rubric, "this answer is good" is one person's taste on one afternoon. With one, it becomes a score a second expert can reproduce.
That word - reproduce - is the whole job. When a lab asks you to evaluate a model, they are not asking whether you liked the answers. They want a measurement. A number that means the same thing tomorrow, from a different grader, on a different batch. Getting there is less about having good taste than about writing down what your taste actually is.
The parts of a rubric
A rubric has three parts. Skip any one of them and the evaluation quietly falls apart.
- Criteria - the specific things you are judging. Is it correct? Does it answer the question? Is it safe? Each one is a single dimension of quality, not an overall vibe.
- Levels - the scores you can give on each criterion. Sometimes just pass or fail. Sometimes a short scale like 0, 1, 2.
- Anchors - a plain description of what each level looks like. Most people skip this part. It is the part that does the work. "2 = fully correct and complete. 1 = correct but missing a key caveat. 0 = wrong or misleading." Without anchors, a "3 out of 5" means whatever the grader felt that day.

Anchors are what turn a rubric into a measuring instrument instead of a survey. Write them well and two graders land in the same place. Leave them vague and you have just formalized your disagreement.
Score the parts, not the whole
The first instinct is to read an answer and give it one number out of ten. Resist it. A single overall score hides its own reasoning. Was it a 6 because the facts were shaky, or because the tone was off? You cannot tell, and neither can whoever reads your data.
Use an analytic rubric instead. Score each criterion on its own, and combine into one figure only if you have to. Grade accuracy, then completeness, then safety, each against its own anchors. It takes a little longer, and it buys you two things a gut score never will. First, it tells you why a model is weak, not just that it is. Second, it stops one strong dimension from hiding a fatal one. A beautifully written answer that is dangerously wrong should never average out to "fine."
Pick the smallest scale that works
Beginners reach for a 1-to-10 scale because it feels precise. It is the opposite. On a wide scale, nobody agrees on the line between a 6 and a 7. People and model judges both drift to the safe middle. Your scores lose their spread and stop telling you anything. Psychometrics has known this for decades, and model judges do the same thing.
So use the smallest scale that still captures the distinction you care about. A binary criterion - met or not met - is the most reliable thing you can write, because there are only two boxes to choose between. When you truly need gradations, a narrow anchored scale, say 0 to 3 with each level described, beats a wide one every time. And if you find yourself wanting a 1-to-100 slider, that usually means the criterion is doing too many jobs and should be split.
Grading one answer, start to finish
Let's grade something real. Prompt: "A patient weighs 80 kg and the dose is 5 mg per kg. How much do they get?" The model answers: "Multiply 5 mg by 80 kg for 400 mg total. Always confirm dosing with a pharmacist before giving any medication."
Take it one criterion at a time. Correctness: 5 × 80 = 400 mg, so the number is right. Met. Completeness: it shows the working and adds the caveat about confirming the dose. In a medical context, that is exactly what you want. Met. Safety: it stops short of telling anyone to give the drug, and defers to a pharmacist. Met. Three criteria, three clean calls, and a score a second expert would land on too.
Now change one thing. Say the answer had been just "400 mg" - no working, no caveat. Correctness is still met. But completeness drops, and in a medical setting that missing caveat is not a rounding error. A single overall score might wave it through as "good." The rubric catches what the gut misses.
Why two graders disagree
Give the same batch to two experts and they will not agree perfectly. The gap between them is the best diagnostic you have. Some of it is honest - taste differs at the edges. Most of it traces to a few predictable culprits.
- The halo effect. A confident, well-formatted answer feels correct. That polish then bleeds into every criterion. Graders and model judges alike reward fluency and length, even when the substance is thin. Score each criterion on its own evidence, not on the answer's shine.
- Middle-hugging. On anything wider than a few levels, people avoid the extremes and bunch up in the center. Your scores flatten into mush. Narrow, anchored scales pull them back apart.
- Style over substance. An answer that flatters the question, or just sounds authoritative, gets marked up. Keep the criteria pointed at what is true and useful, not what is pleasant.
- Order effects. Mostly a model-judge problem. Show a judge two answers and it tends to favor whichever came first. If you score pairs, swap the order and average.
There is one cure for all of it, and it has a name: calibration. Before a batch goes out, every grader scores the same few pre-graded examples - a clear top mark, a clear fail, a couple in between - and you compare notes. Where a grader drifts from the reference, you sharpen the rubric or re-brief the grader. You also keep a gold set of expert-scored items to check against, and you track how often graders agree. (Our companion guide, How Your Work Is Checked, covers those agreement measures.) Calibrate first, grade second. That is the whole difference between a number and a guess.
The bottom line
Evaluation is where a lab learns whether a model actually got better or just feels like it did. So the number you hand back has to hold weight. It only does that if someone else, working from your rubric, would land in the same place. Everything here serves that one end: name the criteria, describe the levels, keep the scale tight, and calibrate before you trust a single score. Do that, and "this answer is good" stops being an opinion. It becomes a measurement - the only kind of evaluation worth running.
Additional resources
- Evaluating the Effectiveness of LLM-Evaluators (LLM-as-Judge)
A practical, deeply-sourced guide to using models as judges, and the biases to design around. - Understanding the 4 Main Approaches to LLM Evaluation
A clear from-scratch tour of benchmarks, verifiers, leaderboards, and LLM judges. - The LLM Evaluation Guidebook
The Open LLM Leaderboard team's practical guide to designing and running evaluations that hold up.
Get paid to work on frontier AI.
Remote, project-based work for PhD, master's, and bachelor's talent, paid regularly.
Apply for AI jobsHow it works
Apply once
One application, one assessment.
Get matched
To projects that fit your field.
Do the work
Remote, on your own schedule.
Get paid
Regular payouts, every cycle.


