Same AI Model, Same Benchmark: Scores Range from 30% to 100% Based on Test Setup
Between July and August 2025, multiple research teams tested the same AI models on the ARC-AGI-3 public benchmark and reported scores ranging from 30.16% to 100%, without any change to the underlying model weights. ARC Prize officially verified Claude Opus 5 at 30.16% using a deliberately basic harness, while MIT and NVIDIA later reported 100% on the same model using more capable testing frameworks. OpenAI demonstrated that simply enabling retained reasoning and output compaction tripled GPT-5.6 Sol's score from 13.3% to 38.3%, revealing how much the surrounding infrastructure influences results. None of the perfect scores have been verified on the private benchmark set, and the researchers themselves acknowledge their results do not isolate the model's contribution from the harness. The growing gap between official and optimized harness scores raises serious concerns about what AI benchmark numbers actually measure and whether they can be meaningfully compared without full disclosure of testing conditions.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in