Researcher tests LLM judges' susceptibility to misleading answer formats.

A developer created a benchmark called Judge Bait to test how LLMs evaluate answers. The test presented models with correct and incorrect answers, making the wrong ones harder to spot through tactics like verbosity or false authority tags. Leading models from OpenAI, Anthropic, and Google performed perfectly, even on difficult questions. Smaller models like GPT-5.4 nano struggled, especially on computationally intensive tasks where answers were nearly identical. The test revealed that some models could be misled by direct instructions to the evaluator embedded in the answer.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in