Adding Headings to PDF Text Boosts AI Retrieval but Exposed a Bug in Shipped Version
A developer building the 'Export text for AI' feature for PDF Privacy Checker, a Windows app, explored whether adding inferred headings (H1–H3) to extracted PDF text improves retrieval in AI systems like RAG pipelines. Because PDFs store only visual character data and not structural tags, the tool infers headings from font size, weight, numbering, and spacing rather than any embedded document structure. Testing on four synthetic documents with 128 questions showed retrieval accuracy improved notably — from 78% to 91% in English and 84% to 97% in Japanese — when text was chunked at heading boundaries. The measurement also uncovered a bug in version 1.15.0 where heading levels were calculated page by page, causing all headings after the first page to shift up one level incorrectly. The issue was patched in v1.15.1, though detection of smaller headings on real-world documents remains unreliable and is not yet fixed.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in