AI Code Review Benchmarks Are Only as Reliable as Their Bug-Counting Methods
A September 2026 benchmark by Entelligence tested two AI code review models, GPT-5.6 Luna and GPT-6 Astra, against 50 public pull requests from well-known open-source projects. Astra detected more bugs with higher precision (96%) but cost nearly 28 times more than Luna, which still achieved 74% precision. The benchmark stood out for publishing its methodology, including a dual-judge verification system where both reviewers had to agree before a finding counted as a real bug. Experts note that most AI tool comparisons omit such details, making precision and recall figures difficult to trust without knowing who labeled the data and how disagreements were resolved. Teams evaluating code review tools are advised to demand transparency on labeling protocols, judge independence, and inter-rater disagreement rates before acting on any benchmark claims.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in