AI Crawlers Quietly Drain Server Resources, Posing Hidden Costs for the Web
On August 29, kernel.org administrator Konstantin Ryabitsev revealed that AI scrapers were consuming 14–16 of the site's 90 CPU cores continuously, just to render individual Git HTML pages that were already available for bulk cloning. The redundant requests highlighted how automated retrieval can impose real infrastructure costs on source websites without disrupting them visibly. AI-related web traffic spans distinct activities — model training, search indexing, and user-triggered fetches — each carrying different implications for site owners and developers. Agentic workflows can worsen the problem when multiple subagents independently fetch the same content, adding latency and server load without improving the quality of results. Experts argue that AI application developers need to optimise for distinct, relevant evidence rather than raw page volume, and should consider the server-side cost their retrieval behaviour imposes on publishers.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in