VIDRAFT's AX-RAY Framework Detects Hidden Causal Leakage in AI Language Models
Korean AI safety startup VIDRAFT published a research paper on arXiv on August 24, 2026, introducing a diagnostic framework called AX-RAY designed to detect causal leakage in autoregressive language models. Causal leakage occurs when future token information illegitimately influences earlier positions in a model, undermining the correctness guarantees that such models rely on. The method works by running two nearly identical inputs through a model and comparing internal activations layer by layer, requiring no gradient computation or retraining. When tested on public models, the framework identified leakage in Nemotron-H-8B and Zamba2-1.2B, and successfully caught all 192 synthetic faults injected during validation. VIDRAFT has filed a Korean patent on the underlying technology and is positioning AX-RAY as a verification tool for government-backed AI security programs.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in