Experiment Shows AI Models Default to Negative Verdicts Without Real-World Evidence
A developer ran a controlled blind experiment across five AI model families — Grok, DeepSeek, GPT, Gemini, and Claude — asking each to argue whether AI makes humanity intellectually stronger or weaker. All ten runs, across two isolated arms each, unanimously concluded 'weaker,' drawing on the same well-documented research around cognitive offloading and memory decline. To investigate the cause, the developer ran a follow-up test with Gemini using four conditions, including one where the model was given a real three-month record of an individual's AI-assisted work, complete with failures and no instructed conclusion. That single change — introducing actual evidence — flipped the verdict to 'stronger' in both draws, while reworded or reframed questions without evidence still returned negative results. Analysis of the models' reasoning traces revealed they had defaulted to the 'weaker' position largely because the atrophy-focused literature is older, larger, and more citable, exposing how AI models can mistake bibliographic abundance for truth.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in