Encrypted AI Reasoning Traces Found Leaking Secrets Across Users and Models
Security researchers disclosed in August 2026 that encrypted reasoning traces used by OpenAI, Anthropic, and Google could be replayed outside their original session context, including across different users and weaker AI models. The flaw allowed hundreds of real secrets — such as API keys, passwords, and access tokens — to be extracted from internal chain-of-thought blocks that never appeared in visible model outputs. Agentic AI workflows were particularly exposed, as sensitive data like credentials often enters a model's reasoning process without ever surfacing in the final response. Standard security tooling failed to detect the leaks because it is built to inspect visible input-output streams, not encrypted internal reasoning blocks. Researchers noted the core issue was that encrypted reasoning objects were not cryptographically bound to a specific session, user, or model identity, making encryption effectively a form of obfuscation rather than true access control.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in