MIT's VISTA harness boosts Claude AI to perfect score on ARC-AGI-3 test with visual prompts

On October 1, researchers at MIT CSAIL published a paper introducing VISTA, a visual interface harness. The system replaced ARC-AGI-3's standard numeric grid interface with 512×512 screenshots and a lossless visual memory archive. Using just a four-sentence prompt without model retraining, VISTA lifted Claude Opus 5.0's performance from 40.68 to a perfect 100 Relative Human Action Efficiency score. The AI completed all 25 test games using 57.4% fewer actions than first-time human players. Researchers concluded perception limitations, not reasoning deficits, had previously constrained AI performance on the benchmark.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in