Over Half of AI Crawler Traffic on One Site Could Not Be Verified as Genuine
A web publisher analyzing 13,491 crawler requests to their site over 30 days found that more than half came from AI vendors — including Meta, ByteDance, and Amazon — that publish no IP ranges or verification method, making it impossible to confirm the requests' true origin. A further 2,226 requests could not be verified because the publisher's local copy of vendor IP lists was outdated, while 4,318 requests came from addresses that did not match any published list. Notably, all 984 requests claiming to be Perplexity-User and all 103 claiming to be PerplexityBot had a 0% verification rate, while Applebot and GoogleOther scored 98% and 100% respectively. The analysis highlights that a user-agent string is simply a self-reported claim — anyone can spoof it with a single command — making unverified crawler statistics potentially unreliable. The findings point to a structural gap in AI crawler accountability, where site owners currently have no reliable way to confirm the identity of the majority of automated traffic they receive.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in