Medium blocks major AI training crawlers but remains open to AI search indexing
A developer discovered that Medium's robots.txt file explicitly blocks eight AI-related crawlers — including GPTBot, ClaudeBot, and Bytespider — preventing bulk data collection for model training. However, the platform remains accessible to search-grounded retrieval bots such as Googlebot, Bingbot, PerplexityBot, and OAI-SearchBot, meaning Medium articles can still appear as cited sources in AI-powered search tools. Notably, Google-Extended and CCBot are also left unblocked, leaving open indirect pathways to AI training through Google's existing crawl data and the Common Crawl archive. The finding prompted the developer to abandon plans for a Medium column on AI search visibility, since the platform's setup works against the strategy being advised. A comparison of similar platforms — including DEV, Substack, and Hacker News — shows none have singled out AI crawlers the way Medium has, while Reddit takes the opposite approach by blocking all crawlers and licensing its archive directly.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in