LLM as a judge
LLM as a judge means using one model to score or label the answers of another · fast and cheap evaluation for answers that code cannot grade exactly.
Text reviewed October 5, 2026
LLM as a judge means using one model to score or label the answers of another · fast and cheap evaluation for answers that code cannot grade exactly.
Basic
Some answers can be graded by code: a number, a date, a function that passes its tests. Open-ended answers cannot. A judge model reads the question and the answer, applies a written rubric and returns a label or score.
Deep
Judges are cheaper and faster than human review but have known biases, such as preferring longer answers or answers in their own style. Good practice is a fixed rubric, a judge that sees only what it needs, and human spot checks. BenchGecko uses a judge model with a fixed rubric for the refusal labels in the Censorship Index and the Model Drift Index, and publishes every raw answer.
Expert
Changing the judge model or the rubric changes the results, so they should be versioned like the test itself. Agreement with human labels on a sample is the usual quality check.
Depending on why you're here
- ·An AI that grades other AIs
- ·Use code grading where an exact answer exists
- ·Spot check judge labels
- ·Cheap evaluation makes continuous testing affordable
- ·Version the judge and rubric with the test
- ·Check agreement with human labels