Hard Negative Mining: How 'Almost Right' Examples Make AI Models Smarter
Hard negative mining is a machine learning technique that improves AI models by training them on examples that are nearly correct rather than obviously wrong, forcing finer distinctions. A model learns little from clearly unrelated examples, but struggles—and thus learns more—when presented with closely related alternatives that differ in subtle but important ways. The approach originated in computer vision, notably with Google's FaceNet in 2015, which used difficult face-pair examples to sharpen facial recognition accuracy. The same principle later proved critical in dense text retrieval, with the 2020 Dense Passage Retrieval paper demonstrating its value for question-answering systems. Today, hard negative mining is widely applied across large language model workflows including RAG pipelines, semantic search, reranking, and embedding training, where the core challenge is semantic confusion rather than outright irrelevance.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.



Discussion (0)
Log in to join the discussion and vote.
Log in