GPT-5.6 Sol jumps from 13.3% to 38.3% on ARC-AGI-3 with config changes alone

OpenAI tested GPT-5.6 Sol on the ARC-AGI-3 benchmark under two different API configurations without changing the model itself, and the score jumped from 13.3% to 38.3% in Relative Human Action Efficiency while using one-sixth the output tokens. The key differences were retaining the model's reasoning across turns and using history compaction instead of truncating older context once conversations exceeded 175,000 characters. The default harness discarded private reasoning after every move and dropped earlier actions as context filled, effectively forcing the model to re-derive its understanding of game rules from scratch each turn. With the adjusted configuration, GPT-5.6 Sol solved all six levels of a game where no frontier model had previously passed the first level on the public leaderboard. The findings suggest that a significant portion of the performance gap between frontier models and human testers may stem from agent loop design rather than model capability.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in