Same fabricated RAG answer scored 0.0 by RAGAS and 1.0 by DeepEval in faithfulness test
A developer building AI tools for healthcare discovered that two widely used LLM-based faithfulness metrics, RAGAS and DeepEval, produced opposite scores when evaluating the same fabricated output from a RAG system. In a controlled experiment using 27 outputs and five repeated runs each, RAGAS correctly flagged the hallucination while DeepEval awarded a perfect score, even praising the answer for accuracy. To address reliability concerns in high-stakes medical applications, the developer built OpenGATE, a deterministic evaluation framework that checks for required facts, numeric traceability, and appropriate abstention without relying on a grader model. Comparative testing showed LLM judges hold a clear advantage in detecting meaning inversions, while deterministic checks proved more consistent for catching omissions and fabricated numbers. The author concludes that neither approach alone is sufficient, recommending a hybrid evaluation strategy combining LLM judges with deterministic checks for production-grade AI systems.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in