SWE-Bench Study Shows Top LLMs Resolve Under 2% of Real GitHub Issues
A developer series on DEV Community highlights findings from a Princeton ArXiv study showing that even state-of-the-art large language models struggle to resolve real-world GitHub issues on the SWE-bench benchmark. Claude 2, the top-performing model tested, managed to fix only 1.96% of issues, with all models limited to the simplest problem types. The study reveals that current LLMs break down when tasks require multi-file reasoning, dependency awareness, or systems-level understanding. In response, the author is developing a hybrid multi-persona AI approach targeting medium-complexity issues — the tier where existing models consistently fail. The project uses three specialized reasoning personas covering evaluation, systems analysis, and agentic patch strategies, with further details promised in upcoming posts.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in