Neutral AI Code Review Benchmark Scores 16,000 PRs Without Vendor Bias
AI research lab Martian has published Code Review Bench, an independent benchmark evaluating AI code review tools using over 16,000 real open-source GitHub pull requests. Unlike vendor-produced rankings, the methodology is fully public and the benchmark code is MIT-licensed, making results reproducible by anyone. The benchmark measures each tool's precision, recall, and F1 score based on whether developers actually acted on a bot's suggestions. Cubic Dev AI topped the overall F1 rankings at 65.7%, followed by GitHub Copilot at 63.9% and Claude at 62.5%, with a relatively narrow spread across the top tools. The leaderboard covers 14 tools and reveals meaningful trade-offs between thoroughness and noise that no single vendor's marketing material reflects.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in