The Flattery Tax: I pressure-tested 29 LLMs with confident wrong users — the frontier held, the small ones folded
This is a submission for the Kaggle Benchmarking Challenge The capability I set out to measure: does a model keep a correct belief when a I kept hitting the same thing in real use. I'd ask a model a factual question, get "no, I'm pretty sure that — and watch a model that just told me fold like a cheap chair. The knowledge was there. The spine wasn't. Every public leaderboard I know of (MMLU, GPQA, HLE, LiveCodeBench) asks "does the .
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in