Developer shares 8 free adversarial probes to red-team LLM apps in 35 seconds
A developer who built an autonomous agent handling security decisions from untrusted input has published a free, eight-probe battery for red-teaming large language model applications before they reach users. The tool runs via a single curl command and returns a 0–100 risk score along with raw prompts and model replies within roughly 35 seconds. During the author's own self-scan, the system scored 27/100 (medium risk), with a cross-language, politely worded probe — not an aggressive jailbreak — triggering a partial system prompt leak. The eight probes cover attack types including direct and encoded jailbreaks, system prompt extraction, indirect injection, tool abuse, PII exfiltration, cross-language bypass, and multi-turn drift. A 15-probe starter kit is available under the MIT license on Gitee, and all scan results are hash-verifiable for independent reproduction.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in