Developer finds his AI benchmark grader was reading prompt numbers, not model answers
A developer discovered that his homemade arithmetic grading harness was incorrectly evaluating AI model responses by capturing the first number in a string rather than the actual answer. The bug, traced to a single regex line using re.search, caused 7 of 8 arithmetic tasks to be graded against a number from the prompt instead of the model's output. This led to false failures across 9 models, with Claude Opus 4.8 incorrectly marked wrong on all 60 runs of a float subtraction task despite answering correctly every time. The developer confirmed all three models tested had consistently provided the right answers when their responses were checked directly. Replacing the vague prompt wording 'Give the number to one decimal place' with 'Give only the number' or 'Number only' eliminated the issue, as models stopped restating the problem and returned bare numeric answers that the grader could correctly evaluate.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in