Training-Free Head Pruning Cuts 20% of Attention FLOPs in Diffusion Transformers
Researchers have found that structural template tokens in diffusion transformers act as implicit semantic registers, absorbing text semantics while allowing a significant portion of attention heads to be safely removed. A training-free pruning rule that targets heads attending most strongly to prompt tokens eliminates roughly 20% of joint-attention FLOPs with only a 1.4-point drop in GenEval scores. Around 20–30% of attention heads can be pruned without retraining, additional data, or gradient-based saliency analysis. The finding marks a departure from prior pruning strategies, which treated all heads as equally essential and relied on weight magnitude or costly fine-tuning loops. The study is currently limited to text-to-image diffusion transformers evaluated on GenEval, and its applicability to other architectures or quality metrics remains to be tested.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in