Researchers Expose $720 Attack That Decrypts Hidden AI Reasoning Across Major LLMs

A research paper published on August 11, 2026, revealed a series of attacks against the encrypted reasoning systems used by OpenAI, Anthropic, and Google. Exploiting a key management flaw common to all three providers, researchers showed that encrypted chain-of-thought blocks could be decrypted by using a weaker sibling model as an unintended decryption oracle. The entire process of decoding 10,000 reasoning traces cost approximately $720, enabling potential IP theft, credential harvesting, safety filter bypasses, and persistent prompt injection. Over 315,000 encrypted reasoning blocks had already been scraped from public repositories on GitHub and Hugging Face before the paper's release. All three AI providers reportedly patched the core vulnerability before publication, but the findings have raised broad security concerns for developers building on LLM APIs.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in