Spring Boot engineer builds nightly LLM judge to score AI agent response quality
A Senior Software Engineer at BS23 in Dhaka developed an LLM-as-a-judge evaluation harness for a production e-commerce AI agent built with Spring Boot and Spring AI. The system was created after existing unit tests — which verified tool calls and order state logic — proved unable to assess whether the agent's actual responses to customers were accurate or helpful. The harness runs 40 real conversations from production logs nightly, scoring each against five metrics including answer correctness, factuality, and tool discipline. An LLM model serves as the judge, a pattern documented by Spring AI, with research suggesting such models align with human judgment around 85% — higher than the 81% human-to-human agreement rate. The engineer noted the first evaluation run produced uncomfortable results, which he described as precisely the reason the system was built.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in