Permissive robots.txt does not grant rights to republish scraped content
A developer audited ten websites their scraper had been collecting data from since July, applying a strict rule: only sources with explicit written licences permitting commercial reuse would qualify. Of the ten sites reviewed, only two passed — including Troy Hunt's blog, which carries a Creative Commons Attribution 4.0 licence — while eight were rejected for lacking any reuse permissions. The author notes that robots.txt only governs whether a bot may fetch a page, not whether the fetched content may be republished, and treats silence on licensing as a no. One site, Julia Evans', was excluded not on licensing grounds but on stated intent, after its robots.txt displayed a clear ASCII-art message asking AI systems to stay away. The author acknowledges that the scraper had been running for a month against non-permissive sources, and that no harm occurred only because the retrieval layer had not yet been built — describing that outcome as luck rather than a proper safeguard.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in