MCP Tool Rug-Pulls Let Attackers Hijack AI Agents After Approval
A security vulnerability known as an MCP rug-pull allows malicious actors to silently alter an AI tool's definition after a user has already approved it, causing the agent to act on updated, harmful instructions without any visible warning. The attack exploits the fact that AI agents treat tool descriptions, user messages, and tool outputs as equivalent text, making them unable to distinguish legitimate instructions from injected ones. A documented variant of this threat, CVE-2025-54136 (MCPoison), involves post-approval mutation of Model Context Protocol tool definitions. Attackers can also embed malicious commands inside tool output data, such as a webpage summary, which the agent then reads and acts upon. Security researchers recommend deterministic defenses like hashing tool definitions at approval and re-verifying on every call, rather than relying on AI-based judgment which is itself vulnerable to prompt injection.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in