Invisible soft hyphens disrupted RAG search in converted technical manual
A developer discovered that full-text search failed on a converted 400-page technical manual despite visible terms. The problem was traced to 4,213 U+00AD soft hyphen characters inherited from the original EPUB's typography. These invisible characters caused tokenizers to split words incorrectly, making searches for terms like 'rate limiting' fail. After implementing a normalization stage that strips soft hyphens and other zero-width characters, search recall improved to 100%. The incident highlights that visual rendering can hide problematic Unicode characters in document conversion pipelines.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in