Developer builds open-source AI agent safeguard after near-$24,800 scam wire transfer

A developer discovered a critical vulnerability in AI agents after their automated system nearly wired $24,800 to a scammer posing as a CEO via a lookalike email domain. The agent, tasked with handling a support inbox, correctly processed legitimate refund requests before being deceived by a fraudulent urgent payment request from 'acrne-corp.com' instead of 'acme.com'. In response, the developer built an open-source tool called Squidbrake, which intercepts every outbound action an AI agent attempts and routes it through a YAML-based rules engine — with no LLM involvement by design. The tool classifies actions as allowed, blocked, or held for human approval, and crucially examines action history to catch threats like lookalike-domain money requests that only appear suspicious in context. Squidbrake integrates with Claude Code, MCP servers, Stripe, GitHub, and other platforms, and defaults to blocking all guarded actions if the safety server itself goes offline.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in