LLM Judges Carry Systematic Biases That Can Skew AI Evaluation Results
A 2023 NeurIPS paper by Lianmin Zheng and co-authors identified three key biases in LLM-as-a-judge evaluation systems: position bias, verbosity bias, and self-enhancement bias. While strong judges like GPT-4 agree with human preferences over 80% of the time, these systematic biases mean that simply running more evaluations does not correct skewed results. Unlike random error, systematic bias sits in the expected value, so increasing sample size only makes a flawed measurement more precise rather than more accurate. Popular eval frameworks including DeepEval, RAGAS, OpenAI Evals, and Langfuse largely rely on LLM judges without fully addressing these distortions. A smaller set of tools, such as promptfoo and Future AGI's eval SDK, combine deterministic scorers with LLM judges to reduce over-reliance on a single biased instrument.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in