AWS Strands Agent Lifts Claude Opus 5 from 30% to 99.95% on ARC-AGI-3 Benchmark
AWS engineers used Claude Opus 5 paired with the open-source Strands Agents SDK to build an AI agent that completed all 183 levels across ARC-AGI-3's 25 public environments, achieving a 99.95% relative human action efficiency score in a single 8-hour run costing roughly $830 in tokens. ARC-AGI-3, developed by ARC Prize, tests AI systems by dropping them into unfamiliar environments without explaining the rules or objectives, requiring the agent to experiment and adapt autonomously. The same Claude Opus 5 model scored only 30.16% when evaluated by ARC Prize under standard conditions, highlighting how critically the surrounding agent architecture — or 'harness' — influences real-world performance. NVIDIA reported a comparable outcome with its AVO agent, also built on Opus 5, scoring 100% on the same public game set. Developers note that the public benchmark results, while not the official competition evaluation, offer valuable insight into which agent design choices enable AI to handle complex, long-horizon tasks.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in