Same CLI Tool, Wildly Different Scores: Why Harness Names Mislead AI Benchmarks
Specific Labs released Real-SWE, an enterprise-focused software engineering benchmark, revealing a critical flaw in how AI coding tool results are commonly reported. Two models run on the same Codex CLI harness scored 33.8% and 16.2% respectively, a gap of more than 2x despite sharing the same interface. Similarly, two models running on Claude Code differed by 10 percentage points, underscoring that the harness is merely a routing layer, not the reasoning engine. Real-SWE addresses this by framing every result as a model-and-harness combination, a standard most vendors avoid because low scores under their own tooling reflect poorly on them. Experts warn that teams evaluating coding agents should always identify both the underlying model and the harness, as the model swap is the single biggest performance lever yet remains largely invisible in vendor marketing.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in