Hybrid Search Plus LLM Reranking Lifts Knowledge Base Recall to 0.87
A four-part experiment tracked the rebuilding of the Cerebras knowledge base, progressively expanding the document corpus from roughly 3,700 to 16,315 items and the evaluation set from 22 to 31 questions. Pure vector search consistently outperformed hybrid search in top-1 and MRR metrics until a reranking step was introduced in the final phase. Adding an LLM reranker over the fused top-20 hybrid results pushed recall@1 from 0.39 to 0.87 and recall@3 to 0.94, marking the first clear overall improvement. Corpus distillation and comment-bursting raised the recall ceiling at deeper ranks but hurt precision at the top by removing exact error strings from embeddings. The data confirms that hybrid retrieval alone is not a reliable upgrade over dense vector search when query signals are diffuse or answers reside in code chunks.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in