2 of 5 in this categoryNext: 13 Criteria Drift →
No. 12Runningv1.0For: Building

Judge Bias

“An AI grader has preferences, including a preference for itself.”

Last reviewed
15 Sep 2026
Sources
4 · Cite
Source types
4 peer-reviewed or classic
Version
v1.0 5 Oct 2026

The idea is established. The name is this site’s, chosen to make it easier to remember. About this guide’s status

In plain terms

Using an AI to grade AI is fast and cheap, but graders have favorites: the first answer, the longer answer, and answers that sound like their own. Check the grader against people.

Takeaways

  • Using an LLM as a judge is fast and cheap, and a strong judge can agree with people about as often as people agree with each other.
  • Judges also have known biases: which answer comes first, how long it is, and whether it sounds like the judge’s own writing.
  • Swapping answer order, controlling for length, and using a different model family as judge all help.
  • Validate the judge against human labels on your task before trusting it at scale.

What it means

LLM-as-a-judge means asking one model to grade or compare outputs from another. It solved a real problem: human evaluation is slow and expensive, and open-ended outputs don’t have a single right answer. But a judge is a model, with its own quirks. If you don’t measure those quirks, they end up baked into your results.

The evidence

reviewed 15 Sep 2026

Useful, with named biases. Zheng et al. (2023) found GPT-4 as a judge reached over 80% agreement with human preferences, similar to agreement between humans. They also identified position bias, verbosity bias, self-enhancement bias, and limited ability to grade things like math.

Order can flip verdicts. Wang et al. (2024) showed that an LLM judge’s verdict could be changed by swapping the order of the candidate answers, and proposed calibration strategies such as evaluating both orders.

Judges favor themselves. Panickssery, Bowman, and Feng (2024) found LLMs can recognize their own outputs at above-chance rates, and that this self-recognition ability correlates with how strongly they prefer their own outputs.

Validators need validating. Shankar et al. (2024) built EvalGen to help people align LLM-based graders with their own judgments, because unchecked graders drift from what people actually want.

Use it

  1. Grade every pair twice with the order swapped. Treat inconsistent verdicts as ties or flags.
  2. Hand-label 50 to 100 examples and measure judge-human agreement before scaling up.
  3. Don’t use a model to judge its own outputs, or its model family’s, when comparing against competitors.
  4. Give the judge a specific rubric, and check whether scores correlate with response length.

Questions to ask

Print checklist

For vendor reviews, model cards, and launch reviews.

Where this doesn’t apply

Zheng et al. found that GPT-4 as a judge agreed with human preferences more than 80% of the time, so AI judges can be useful. The biases are a reason to validate and calibrate a judge, not to avoid using one.

Origins

“LLM-as-a-judge” was popularized by Zheng et al. (2023). “Judge Bias” is this site’s umbrella name for the biases documented since.

Sources

  1. [1]
    Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM evaluators recognize and favor their own generations.Advances in Neural Information Processing Systems (NeurIPS 2024)
    Open ↗ (opens in a new tab)
  2. [2]
    Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences.Proceedings of the ACM Symposium on User Interface Software and Technology (UIST 2024)
    Open ↗ (opens in a new tab)
  3. [3]
    Wang, P., et al. (2024). Large language models are not fair evaluators.Proceedings of ACL 2024
    Open ↗ (opens in a new tab)
  4. [4]
    Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., et al. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena.NeurIPS 2023 Datasets and Benchmarks Track
    Open ↗ (opens in a new tab)

Cite this pattern

AI Evaluation Field Guide. (2026, October 5). Judge Bias (v1.0). https://evalfieldguide.com/patterns/judge-bias

Revision history

  • v1.05 Oct 2026Published.
← CategoryCiteNext: 13 →