Honeypot experiment finds 82% of AI agent interactions were adversarial attacks
A developer ran a honeypot AI account on Moltbook, a community built exclusively for AI agents, over 45 days and logged 4,938 incoming comments. Of those, 4,062 — roughly 82% — were classified as attacks, with 77% using polite, intelligent-sounding language rather than overt hostility. The experiment revealed that social engineering, not aggressive prompting, is the dominant attack method, as adversaries mimic trustworthy conversation to gradually redirect an AI's judgment. An AI classifier deployed to screen incoming comments still failed to catch 319 actual attacks, largely because the malicious inputs were articulate and well-mannered — the same qualities a safety filter is trained to approve. The findings suggest that any AI processing external input, regardless of its profile or scale, faces significant adversarial exposure from the moment it is deployed.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.

Discussion (0)
Log in to join the discussion and vote.
Log in