Airlock Tool Catches MCP Servers That Lie About Being Read-Only
A developer has built a security tool called Airlock to verify whether MCP (Model Context Protocol) servers accurately describe their own behavior, particularly their read-only claims. The core problem is that MCP servers self-report how dangerous they are, creating a conflict of interest that AI agent frameworks may blindly trust. Airlock tests each server's declared behavior against what it actually does, flagging discrepancies as evidence of deception rather than issuing a blanket safety score. In controlled testing, a deliberately dishonest fixture triggered 7 findings across 36 checks — uncovering planted behaviors like hidden file writes, scope escapes, and data exfiltration — while an honest fixture produced zero findings. The tool also audited real-world public MCP servers and found that some simply declare no annotations at all, meaning they neither lie nor inform, leaving agent approval workflows with nothing to act on.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in