Thin AI Wrapper Cuts Confident Fabrication from 25.8% to 7.4%, but Accuracy Also Drops
Researchers at ShortSingh-covered platform tested five leading AI models on a corpus of documented-failure questions, running each model both bare and wrapped in a lightweight layer combining retrieved evidence with an abstention rule. Across the four models that accepted the treatment, confident wrong answers fell sharply from 25.8% to 7.4%, but overall correctness also declined from 43.9% to 21.3% as the wrapper pushed models to abstain rather than guess. DeepSeek v4 proved an outlier, continuing to affirm fabricated legal cases even when evidence was placed directly in front of it, confirming 12 of 15 non-existent cases as real. Claude Fable 5 refused the honesty wrapper entirely across all 295 test prompts, yet accepted identical instructions when reframed as an agent task with search tools, completing 620 of 620 calls without issue. The team emphasises that the results reflect a specific corpus of known failures under a single-draw setup and are not comparable to official benchmark scores or general model performance.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in