Cheap vs. premium AI code review: why aggregate benchmark scores mislead teams
A comparison published by Entelligence on September 14, 2026, pitted GPT-5.6 Luna ($0.0041 per review) against GPT-6 Astra ($0.113 per review) across the same pull requests on several open-source codebases. While Astra outperformed Luna overall — finding 92 verified bugs at 96% precision versus Luna's 69 at 74% — the gap varied sharply by codebase, with Luna performing comparably on everyday application code but falling significantly short on security-sensitive code like Keycloak. On authentication-related bugs specifically, Luna caught only 9 of 24 versus Astra's 19, and missed multi-step logic flaws involving permission overrides and reusable recovery codes. A repeatability test further revealed that neither model consistently reproduced its own findings across multiple runs, exposing instability that single-run benchmark scores conceal. The authors caution that because surrounding code in the benchmark is publicly available and predates both models' training cutoffs, scores likely overstate real-world performance on unseen codebases.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in