Researchers Expose Side-Channel Flaw That Leaks AI Models' Internal Reasoning
A team of researchers from the University of Tübingen, ETH Zürich, Oxford, and CSET has discovered a previously unknown vulnerability in frontier AI systems that allows attackers to extract hidden chain-of-thought reasoning traces. The flaw exploits the practice of off-loading encrypted reasoning computations to client devices, where a shared decryption key across model families creates a single point of failure. By replaying encrypted traces to smaller, less-aligned variants of the same model family, attackers can recover the original model's internal reasoning in plain text, potentially exposing sensitive data like passwords and API keys. The attack was demonstrated on proprietary models Claude Opus 4.8 and GPT 5.6 Sol, with open-weight model Kimi K3 by Moonshot AI reproducing nearly identical reasoning traces, suggesting possible distillation from closed systems. Other open models including DeepSeek and Inkling showed no such similarity, indicating the vulnerability depends on specific training pipelines and data-sharing practices.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in