Popular AI Agent Prompt Injection Defenses Fail Over Half of Real Attacks, Tests Show
A developer benchmarked ten open-source prompt injection detectors against 629 real-world attacks drawn from an academic dataset, finding that the best detector caught only 51% of attacks at a 2% false-positive rate. The tests revealed that attacks embedded in ordinary content like emails or documents — rather than obvious override phrases — routinely evade word-based detection methods. Meta's Prompt Guard 2 caught just 6 of 629 attacks at its default threshold setting, though recalibrating the threshold dramatically improved its catch rate to around 99%. The researcher notes that even well-tuned detectors function more like smoke alarms than locks, since attackers can retry indefinitely and a single missed injection could trigger harmful actions. The benchmark tool, called buried-injections, is open source and available on GitHub for teams to test their own setups.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in