Attention Maps Show Information Routing, Not Model Reasoning, Research Confirms
Attention maps in transformer models display weight coefficients indicating how much each position attends to others, but do not reveal why a model produces a given output. A 2019 NAACL paper by Jain and Wallace demonstrated that attention weights often disagree with other attribution methods and that alternative attention distributions can yield identical predictions, undermining their explanatory value. A follow-up study by Wiegreffe and Pinter at EMNLP 2019 argued that the claim was too broad, noting that adversarially found distributions are not ones the model naturally learns and that 'explanation' needs clearer definition. The core technical reason attention weights are unreliable explanations is that value vectors across positions tend to be similar after multiple layers, making the weighted sum robust to redistributions of attention. Researchers now broadly treat attention maps as wiring diagrams showing possible information flow within a single layer, not as causal accounts of model behavior.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in