OpenAI harness fix, EU AI gigafactories, and Cognition's SWE-1.7: AI roundup Aug 1, 2026
OpenAI revealed on July 29 that GPT-5.6 Sol's poor ARC-AGI-3 benchmark score of 7.8% was caused by a flawed evaluation harness that discarded reasoning between moves, not by the model itself. Enabling retained reasoning and context compaction in the Responses API lifted Sol's score from 13.3% to 38.3% and allowed it to clear all six benchmark levels, a feat no frontier model achieves on the official leaderboard. The finding highlights that benchmark results reflect the combined effect of model, harness, and settings rather than model capability alone. Meanwhile, the European Commission launched a formal tender on July 30 to build up to seven AI gigafactories across the EU, backed by up to €10 billion in public funding and a €30 billion total target including private investment, with awards expected in early 2027. Separately, AI coding lab Cognition released SWE-1.7 in July, a new frontier coding model built on a Kimi K2.7 base and trained via reinforcement learning, featuring innovations in sampling stability, multi-continent training, and extended task horizons.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in