Study Warns AI Training on AI Output Risks 'Model Collapse' as Web Fills With AI Text
A peer-reviewed 2024 Nature study by Ilia Shumailov and colleagues found that generative AI models progressively lose diversity and accuracy when trained repeatedly on AI-generated content rather than human-created data. The phenomenon, called model collapse, causes rare facts, minority styles, and unusual information to fade first, leaving models that produce confident but bland, narrowed outputs. Researchers described the degradation as occurring in two stages — early loss of rare data and eventual full collapse toward a uniform average — and warned that the damage can be irreversible. The concern has grown more urgent following a Pew Research Center report published on August 20, 2026, which found that over a third of web pages published since ChatGPT's launch show signs of AI authorship. Experts note that while collapse is real and demonstrated, its severity in practice depends on how much unfiltered synthetic data enters training pipelines, making it a matter of degree rather than an immediate catastrophic failure.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in