Developer finds stricter eval rule discards more data as test repetitions grow
A software developer running a frozen 159-task benchmark suite to compare two AI models discovered that a conservative measurement rule increasingly discarded tasks with each additional test repetition. The stricter rule, designed to filter unstable results, threw away 13 tasks across three repetitions and 17 across six, shrinking the informative dataset and eventually masking a statistically significant difference between the models. A second, less restrictive rule was developed and pre-registered before new test runs began, allowing an independent replication on unseen data. The looser rule reached statistical significance on the fresh data, while pooling all six repetitions under the strict rule returned a non-significant result. The developer concludes that conservative stability filters are not cost-free choices, as they systematically discard genuinely stochastic tasks the more data is collected.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in