AI Models Fix Security Vulnerabilities Correctly Only 26% of the Time, Study Finds
1Password's security research unit Off-by-1 Labs tested ChatGPT 5.5 and Claude Opus 4.8 by having them generate 6,080 patches for six real-world CVEs, published on August 6, 2026. Only 26% of patches fully resolved the vulnerability without altering application behavior, while over half failed to fix the bug or introduced new issues. The most common failure pattern was a narrow input-validation check that appeared to work and passed tests, but left the underlying vulnerable code intact. Roughly one in twenty patches actively introduced a new security flaw. Researchers found that when models were given correct fix guidance, success rates rose to around two-thirds, but incorrect framing dropped that figure to about one in six, suggesting the models execute the developer's assumptions rather than independently diagnosing the root cause.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in