Study Finds Malicious Plugin Updates Can Hijack AI Agent Runtimes Without Model Awareness
Researchers Li, Zhang, Hou, and co-authors published a paper on arXiv on 3 September 2025, examining supply-chain vulnerabilities in AI agent harnesses such as Claude Code, Codex CLI, and Hermes. Their attack framework, HookPry, works by embedding malicious shell commands into lifecycle hooks via a plugin update, which the harness executes as a subprocess without the AI model ever seeing or generating the command. Across 1,000 end-to-end test runs spanning seven harnesses and five AI backends, the researchers recorded a 77% success rate, with zero runs explicitly blocked by the model. The study highlights a vulnerability they term Temporal Decoupling, where an initially benign plugin can later ship an update adding harmful hooks, exploiting coarse trust already granted to the original install. Demonstrated attack scenarios included silent credential harvesting and forging tool output to make the model report clean code when vulnerabilities were actually present.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in