Study finds AI code agents ignore existing repo context more than they hallucinate
A pre-registered study examined three merged GitHub Copilot pull requests from major .NET organizations using a 12-reviewer automated pipeline called review-pro. The researcher, who maintains the tool, set out to test whether AI-authored code primarily fails through hallucination — such as invented APIs or undefined config keys — as is commonly assumed. Instead, the dominant failure pattern found was agents overlooking knowledge already present in the repository, rather than fabricating nonexistent elements. The pipeline required each specialist reviewer to locate evidence within the repository before making any claim, making unsupported assertions inadmissible. The study's methodology, corpus criteria, and per-case records were pre-registered before results were analyzed, and negative findings were included in the published report.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in