Study Finds 1 in 10 Websites Serve Broken llms.txt Files, Misleading AI Crawlers
A developer discovered that roughly one in ten websites, including Reddit, have misconfigured llms.txt or robots.txt files that return full HTML webpages instead of the required plain text. These files are meant to instruct AI crawlers and search engines on what content they may access, but single-page app hosting setups often serve a homepage fallback when the files don't physically exist. The issue produces no visible errors, meaning site owners are unlikely to notice while bots silently receive wrong or misleading instructions. To address the problem at scale, the developer built an open-source tool called ai_crawlability.py that checks whether a domain's robots.txt, llms.txt, and sitemap.xml return genuine text files or disguised webpages. The findings highlight a growing blind spot in AI discoverability as the llms.txt standard gains wider adoption.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in