AalborgMBZUAIVILA-LabFeb 19, 2026arXiv:2602.17664

Sink-Aware Pruning for Diffusion Language Models

Aidar Myrzakhan, Aidar Myrzakhan, Tianyi Li, Tianyi Li, Bowei Guo, Bowei Guo, Shengkun Tang, Shengkun Tang, Zhiqiang Shen, Zhiqiang Shen

AI Summary

The paper investigates the role of attention sinks in Diffusion Language Models (DLMs) and finds that, unlike in Autoregressive (AR) LLMs, DLM attention sinks exhibit high variance across denoising steps, suggesting they are less structurally important. Based on this finding, they introduce Sink-Aware Pruning, a novel pruning method that identifies and prunes unstable sinks in DLMs. Experiments demonstrate that Sink-Aware Pruning achieves a superior quality-efficiency trade-off compared to existing pruning techniques without retraining.

Key Contribution

Attention sinks, considered essential in autoregressive language models, turn out to be surprisingly prunable in diffusion language models, leading to better efficiency.

Abstract

Diffusion Language Models (DLMs) incur high inference cost due to iterative denoising, motivating efficient pruning. Existing pruning heuristics largely inherited from autoregressive (AR) LLMs, typically preserve attention sink tokens because AR sinks serve as stable global anchors. We show that this assumption does not hold for DLMs: the attention-sink position exhibits substantially higher variance over the full generation trajectory (measured by how the dominant sink locations shift across timesteps), indicating that sinks are often transient and less structurally essential than in AR models. Based on this observation, we propose ${\bf \texttt{Sink-Aware Pruning}}$, which automatically identifies and prunes unstable sinks in DLMs (prior studies usually keep sinks for AR LLMs). Without retraining, our method achieves a better quality-efficiency trade-off and outperforms strong prior pruning baselines under matched compute. Our code is available at https://github.com/VILA-Lab/Sink-Aware-Pruning.

Architecture Design (Transformers, SSMs, MoE)Inference & Quantization Natural Language Processing

Citation Metrics

Citations0

Influential citations0

References38

Year2026

VenueN/A

Related Papers

Finding related papers...

Search

Sink-Aware Pruning for Diffusion Language Models

Related Papers