Study Finds Small Batch Sizes Skew LLM Rankings in Social Deception Tests
A developer built a benchmarking tool called 'detection lift' to evaluate how well large language models identify liars in simulated Mafia games, aiming for a more precise metric than simple win rates. Across 522 games involving 19 models, early results appeared to show larger models outperforming smaller ones, but those rankings reversed or collapsed entirely as more games were played. A rule-based control agent with no variability scored anywhere from -0.125 to +0.416 across batches, a range wider than the gap between any two AI models tested. The author determined that batches of fewer than 15 games produced all extreme readings, while batches of 30 or more games yielded far narrower and more stable results. The findings are described as exploratory and highlight how small sample sizes can make random noise look like meaningful performance differences when ranking language models.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in