Developer releases open dataset tracking 28 AI crawler user agents and their behavior
A developer has published an open dataset cataloguing 28 AI crawler user agents, including GPTBot, ClaudeBot, and PerplexityBot, available on GitHub and via a live JSON endpoint. The dataset distinguishes between crawlers used for model training and those used for AI-powered search indexing, helping site owners make more targeted blocking decisions. Each entry includes a purpose field and a compliance rating indicating whether the bot actually honors robots.txt directives in practice, not just in vendor documentation. The developer notes that robots.txt alone is insufficient enforcement for bots rated non-compliant, recommending pairing it with WAF or CDN-level rules. The dataset is reviewed monthly, and contributions flagging missing or outdated crawler entries are welcomed.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in