TomeVault Releases 229,375-File AI Instruction Corpus Under Open License
The TomeVault Instruction Corpus (Edition 2026-07) is a newly published dataset containing de-identified structural measurements of over 229,000 AI coding instruction files, including formats such as CLAUDE.md, AGENTS.md, and .cursorrules. The dataset captures one row per retained instruction file instance, meaning a single repository can contribute multiple rows if it contains several such files. No file contents, owner names, repository names, or URLs are included, making the dataset privacy-preserving by design. Each row records structural attributes such as file size, line count, token count, format type, and whether the file references any retired AI model identifiers. Released under a CC BY 4.0 licence with schema version 1.0.0, the corpus is intended to help researchers study how development teams are structuring instructions for AI coding agents.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in