Autonomous AI Agents Now Outperform Human Hackers in Real Bug Bounty Competitions
Autonomous penetration testing agent XBOW reached the top spot on HackerOne's US leaderboard in June 2025, reporting over 1,060 vulnerabilities — including 54 critical findings — in roughly 90 days at an operating cost of around $18 per hour, compared to $60-plus for human experts. Google's Big Sleep agent independently discovered a zero-day SQLite vulnerability that had evaded automated fuzzers for years, marking the first AI-found zero-day in production software. Anthropic's research agent, previewed in April 2026, identified thousands of critical vulnerabilities across major operating systems and browsers, but the company withheld its release, citing the system as too capable for broad deployment. Stanford and CMU's ARTEMIS agent, tested against a real university network of roughly 8,000 hosts, ranked second overall and outperformed nine out of ten certified human penetration testers in a controlled experiment. Despite these advances, benchmarking research presented at ICML 2025 highlighted a significant gap between lab results and real-world performance, with leading models exploiting only 13–25 percent of real production CVEs versus 87 percent in controlled settings.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in