LLMs Miss Critical Security Bugs in Code Reviews, Benchmark Study Finds

A developer built a benchmark called 'Plausible PR' to test whether four leading AI models — Claude Sonnet 5, GPT-5.5, Gemini 3.7 Flash, and DeepSeek-R1 — could detect security vulnerabilities hidden inside seemingly routine code refactors without being prompted to look for them. Each of the 10 test pull requests was designed to appear like a harmless cleanup while actually removing a security safeguard such as a permission check, rate limit, or input boundary. Claude Sonnet 5 performed best with 9 out of 10, while GPT-5.5 scored lowest at 6 out of 10. Notably, all four models failed to flag a bug where switching to a string comparison caused crashes on non-ASCII input instead of returning a clean authentication failure. The findings suggest AI code reviewers reliably catch keyword-triggered issues like SQL injection but struggle with subtle edge-case vulnerabilities that require tracing unusual inputs through logic.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in