Developer audited a 5.5 GB AI dataset by downloading less than 1% of it
A developer discovered a 5.5 GB Chinese astrology AI training dataset on a popular repository and grew suspicious when the sample count of 518,400 matched a perfect nested loop formula rather than real observations. Using HTTP range requests, they downloaded only 48 MB — about 0.8% of the total — by fetching the ZIP central directory and targeted file shards. The audit revealed that the dataset's 781 data shards appeared algorithmically generated rather than empirically collected. More strikingly, a proprietary content library explicitly excluded from the open-source GitHub repo was found bundled inside the release archive, buried nearly 6 GB deep where few would look. The case highlights both a practical technique for sampling large archives cheaply and a cautionary note about assuming release contents match what a repository publicly discloses.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in