Benchmark flaw found: token limit was silently failing models, not measuring them
Developers behind OmnisBench, an open LLM routing benchmark, discovered a critical flaw after community commenters questioned whether test datasets like HumanEval and GSM8K were too old and likely memorised by models. When the team built a fresh split using only post-training-cutoff problems from LiveCodeBench, scores dropped sharply — but the cause turned out not to be model weakness. A 4,096 output token cap was cutting off reasoning models mid-thought on hard problems, producing blank answers that were graded as failures. The team only caught the error because they publish all raw model responses, allowing them to inspect the empty outputs directly. After raising the token budget for harder tasks, results shifted meaningfully, with the cheapest model scoring 60% and ideal routing recovering performance to 93.3% on fresh problems.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in