ConceptsReading · ~3 min · 68 words deep

LLM as a judge

LLM as a judge means using one model to score or label the answers of another · fast and cheap evaluation for answers that code cannot grade exactly.

Text reviewed October 5, 2026

TL;DR

LLM as a judge means using one model to score or label the answers of another · fast and cheap evaluation for answers that code cannot grade exactly.

Level 1

Some answers can be graded by code: a number, a date, a function that passes its tests. Open-ended answers cannot. A judge model reads the question and the answer, applies a written rubric and returns a label or score.

Level 2

Judges are cheaper and faster than human review but have known biases, such as preferring longer answers or answers in their own style. Good practice is a fixed rubric, a judge that sees only what it needs, and human spot checks. BenchGecko uses a judge model with a fixed rubric for the refusal labels in the Censorship Index and the Model Drift Index, and publishes every raw answer.

Level 3

Changing the judge model or the rubric changes the results, so they should be versioned like the test itself. Agreement with human labels on a sample is the usual quality check.

The takeaway for you
If you are a
Curious · Normie
  • ·An AI that grades other AIs
If you are a
Builder
  • ·Use code grading where an exact answer exists
  • ·Spot check judge labels
If you are a
Investor
  • ·Cheap evaluation makes continuous testing affordable
If you are a
Researcher
  • ·Version the judge and rubric with the test
  • ·Check agreement with human labels
Within limits. Use a fixed rubric, keep the judge blind to which model answered, and check a sample against human labels.