Anthropic's Jacobian Lens Tests Whether AI Signals Actually Drive Model Outputs
Researchers at Anthropic have developed a tool called the Jacobian lens to address a long-standing gap in AI interpretability: the difference between a signal being detectable inside a model and that signal actually influencing the model's output. Unlike earlier methods such as the logit lens or tuned lens, which reveal what a layer's state resembles, the Jacobian lens measures whether internal representations do causal work by computing the model's averaged Jacobian from each layer to the final output. The tool includes four operations — READ, WRITE, PATCH, and ABLATE — that together allow researchers to edit internal quantities and observe selective behavioral changes, such as swapping a concept pattern to shift a model's arithmetic answer. Ablating these so-called J-space patterns caused multi-step reasoning to collapse to near zero, though the model retained fluency, sentiment classification, and fact retrieval, indicating the effect is task-specific rather than general. The work highlights that interpretability claims require causal intervention, not just correlation, to be considered reliable.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.


Discussion (0)
Log in to join the discussion and vote.
Log in