New SWE-Bench ProMax Shows Top AI Models Solve Only 41.2% of Real Coding Tasks

A new benchmark called SWE-Bench ProMax, published on arXiv, challenges the near-90% scores AI companies have been reporting by testing models on real-world refactoring tasks spanning multiple files and programming languages. The study found that the widely used SWE-bench Verified benchmark was unreliable due to three compounding flaws: nearly 60% of unsolved problems had faulty tests, models could reproduce gold-patch solutions memorized from public GitHub training data, and 86% of tasks involved only a single file. SWE-Bench ProMax replaced these with 170 carefully vetted problems across seven languages, averaging 11.4 files and 261.6 lines of changes per task. Under these stricter conditions, GPT-5.2 led all models with just 41.2%, followed by Claude Sonnet 4.6 at 38.8%, while open-weight models GLM-5 and Qwen3.5 each scored 36.5% at roughly one-twentieth the cost. Researchers concluded that open-weight models are closing the gap with frontier proprietary models at a fraction of the price, and that cross-file coordination remains the primary failure mode for all tested systems.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in