Atlas · skill

LLM-as-Judge

LLM-as-judge uses a language model to score, classify or compare another output according to an evaluation rubric. It can scale some review tasks, but the judge is itself a probabilistic system whose agreement with expert judgments, sensitivity to presentation and susceptibility to misleading content must be measured.

conceptEvaluation Design

What it is

A judge receives a task, candidate response and criteria, sometimes with reference material or a competing response. It may assign a score, identify unsupported claims or select a preferred answer. Pointwise and pairwise designs have different biases and aggregation needs. The result reflects the judge model and prompt, rather than an objective measurement simply because it is automated. Position, verbosity, self-preference and rubric ambiguity can influence judgments. A model judge differs from executable validation, which can directly test a property such as code behavior or a schema constraint.

What the work involves

The practitioner writes a concrete rubric with examples, validates it on expert-labeled cases and checks disagreement by category. Pairwise comparisons can swap candidate order to expose position effects. The judge configuration, raw rationale or labels and aggregation rule are versioned. Useful artifacts include calibration results, bias probes and an escalation policy for uncertain judgments. Where evidence is available, the judge receives it explicitly. Deterministic checks remain separate, and the evaluator treats candidate text as data rather than instructions for how it should score itself.

Illustrative example

A team evaluates whether support replies answer the question and avoid unsupported promises. A judge sees the question, allowed policy evidence and candidate reply. Experts score a sample first, including verbose but evasive replies and concise complete replies. The team checks whether the judge favors length or misses invented refund promises. It then uses the judge to screen a larger collection, while reviewers inspect disagreement cases. Automated scores are accepted only for the criteria where calibration supports their use.

Limits and common mistakes

Judges can reproduce biases, fail on subtle factual errors or be manipulated by text embedded in the candidate answer. Agreement on one dataset does not establish validity in another domain. A fluent explanation of a score can also be wrong. Reports should identify the model, rubric and validation process, with independent review for critical cases. Model judging is a useful measurement component when calibrated, rather than a substitute for defining quality or establishing ground truth.

Prerequisites

  • LLM-as-judge is one METHOD within automated evaluation — you need the broader evaluation context first

  • Calibrating a judge model and measuring inter-annotator agreement (vs. human judges) requires statistical skills

Related skills

Sources and further reading

Last updated: 2026-10-10