Synthetic Data Fills AI Training Gap, But Experts Warn Against Overreliance
AI developers are hitting a 'data wall' as high-quality human-generated text for training frontier models is projected to be largely exhausted between 2026 and 2032, according to Epoch AI research. Publishers are increasingly blocking AI crawlers and licensing costs are rising, pushing teams to turn to synthetic data, which already accounts for 30–60% of tokens in some frontier model runs. Gartner forecasts synthetic data will comprise roughly 75% of all AI development data by end of 2026, up from just 1% in 2021. However, experts consistently warn that training exclusively on synthetic data risks 'model collapse,' where output diversity shrinks and errors compound over generations. The recommended approach is a hybrid model — typically 30–50% synthetic data anchored by a curated set of real human-generated examples — and never evaluating model performance solely on synthetic holdout sets.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in