GPT-6 Astra Scores 99.9% on ARC-AGI-3, But Harness Choice Explains the Gap
On 3 September 2026, ARC Prize published official ARC-AGI-3 benchmark results for GPT-6 Astra, showing two distinct scores depending on which evaluation harness was used. Under OpenAI's Provider Adapter at high reasoning, Astra scored 99.9% at a cost of $18,817, while the shared Standard harness at maximum reasoning yielded only 62.7% at a higher cost of $26,098. The key difference lies in memory handling: the Provider Adapter preserves reasoning state and compresses context between requests, allowing the model to reuse prior work, whereas the Standard harness does not. ARC Prize explicitly declined to interpret the result as proof of AGI, with co-founder Mike Knoop stating there is insufficient evidence to make that claim. Both scores remain on the leaderboard with harness labels, and ARC Prize has cautioned that any single unlabeled percentage from this benchmark is not a valid basis for comparison.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in