LLM-as-Judge

LLM-as-judge means using a model to score another model's output against a rubric, instead of checking it by hand or requiring an exact match. It's one of the three layers described on the AI evals page: deterministic checks handle anything with one correct answer, LLM-as-judge handles everything else — whether a free-text answer is actually reasonable, faithful, or well-written — at a scale no person could review manually.

Why it exists

A lot of what needs scoring has no single correct string to match against. "Is this summary faithful to the source," "does this answer actually address the question," "is this tone appropriate" — a person can judge these, but not across thousands of outputs every time a prompt or model changes. A second model call, given the output and a scoring rubric, can do that judgment at a scale a human review process can't keep up with.

Real, documented biases

A judge model isn't a neutral scorer — it has specific, measured tendencies:

  • Position bias. When comparing two responses side by side, whichever one appears first gets favored more often than a fair coin flip would predict — one widely-cited study found the first slot winning 10-15 points more often on the same comparisons, purely from ordering. Swapping the order and averaging the result is the standard fix.
  • Verbosity bias. Longer responses get scored higher fairly consistently, independent of whether the extra length actually added anything — inflating preference for the longer answer by a measured 15-30 points in some studies. Explicitly instructing the judge not to prefer length cuts this roughly in half, though it doesn't eliminate it.
  • Self-preference bias. A judge tends to score outputs from its own model family more favorably than outputs from a different one, even when quality is comparable — tied to how familiar or predictable the judge finds the text it's reading, not to genuine quality.

This isn't a reason to abandon LLM-as-judge — it's the reason the AI evals page's third layer, periodic human calibration, exists. A survey of these methods and their documented failure modes is available from recent research on LLM-based evaluation for anyone building a judge pipeline.

Reducing the impact of these biases

Randomize or swap comparison order and average the result, rather than trusting a single ordering. Put an explicit instruction against favoring length directly in the rubric. Where possible, use a judge from a different model family than the one being evaluated, to reduce self-preference. And check the judge's scores against real human judgment on a sample periodically — a judge that drifts or develops a blind spot only gets caught by comparing it to something outside itself.

When it's the right tool

Reach for it when quality is genuinely subjective or open-ended — summarization, free-text answers, tone, faithfulness to a source — and manual review at the volume needed isn't realistic. Skip it, in favor of a deterministic check, whenever there's actually one correct answer to match against: a wrong format, a wrong category from a fixed list, a malformed structured output. Using a judge for something a simple equality check would settle for free is unnecessary cost and an unnecessary source of the biases above.

In this guide
  1. Why it exists
  2. Real, documented biases
  3. Reducing the impact of these biases
  4. When it's the right tool
  5. FAQ

FAQ

If an LLM judge agrees with a human reviewer most of the time, is it safe to stop checking its work?

No. Agreement measured once doesn't rule out systematic biases like favoring longer answers or a particular ordering — a judge can agree with humans on average while still being reliably wrong in a specific, predictable direction. Periodic spot-checks against real human review are what catch that pattern, not an overall agreement rate.

Does using a stronger, more capable model as the judge fix these biases?

Not on its own. Position bias, verbosity bias, and self-preference bias have been measured across models of varying capability — being a stronger model doesn't automatically make a judge immune to preferring the first option, the longer answer, or its own model family's style.

Practice interview questions on AI evals →