LLM Agreement Tests Can Mislead If Models Are Just Echoing Their Own Priors
A small art project encoding sentences as word-length sequences prompted researchers to test whether large language models could recover intended meaning from length-only data. An initial experiment appeared to show models agreeing far above chance, suggesting meaningful signal recovery. However, a second message reversed the result, with the prior-control arm scoring higher than readings of the actual message. The finding revealed that model agreement was largely driven by shared priors and task framing rather than the encoded content itself. The experiment highlights a critical gap in self-consistency evaluations: without a prior-control baseline, high model agreement can be mistaken for genuine signal.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in