dorsal/arxiv
View SchemaRevealing the Attention Floating Mechanism in Masked Diffusion Models
| Authors | Xin Dai, Pengcheng Huang, Zhenghao Liu, Shuo Wang, Yukun Yan, Chaojun Xiao, Yu Gu, Ge Yu, Maosong Sun |
|---|---|
| Categories | |
| ArXiv ID | 2601.07894vv1 |
| URL | https://arxiv.org/abs/2601.07894 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Masked diffusion models (MDMs), which leverage bidirectional attention and a denoising process, are narrowing the performance gap with autoregressive models (ARMs). However, their internal attention mechanisms remain under-explored. This paper investigates the attention behaviors in MDMs, revealing the phenomenon of Attention Floating. Unlike ARMs, where attention converges to a fixed sink, MDMs exhibit dynamic, dispersed attention anchors that shift across denoising steps and layers. Further analysis reveals its Shallow Structure-Aware, Deep Content-Focused attention mechanism: shallow layers utilize floating tokens to build a global structural framework, while deeper layers allocate more capability toward capturing semantic content. Empirically, this distinctive attention pattern provides a mechanistic explanation for the strong in-context learning capabilities of MDMs, allowing them to double the performance compared to ARMs in knowledge-intensive tasks. All codes and datasets are available at https://github.com/NEUIR/Attention-Floating.
{
"annotation_id": "afb502b4-8ea2-4f3d-ac26-1edb6500a986",
"date_created": "2026-02-17T05:53:12.326000Z",
"date_modified": "2026-02-17T05:53:12.326000Z",
"file_hash": "a77deb7faf7c1ee8491a03f3137bb9fe90d4b8b0546454a3303143836bd16d94",
"private": false,
"record": {
"abstract": "Masked diffusion models (MDMs), which leverage bidirectional attention and a denoising process, are narrowing the performance gap with autoregressive models (ARMs). However, their internal attention mechanisms remain under-explored. This paper investigates the attention behaviors in MDMs, revealing the phenomenon of Attention Floating. Unlike ARMs, where attention converges to a fixed sink, MDMs exhibit dynamic, dispersed attention anchors that shift across denoising steps and layers. Further analysis reveals its Shallow Structure-Aware, Deep Content-Focused attention mechanism: shallow layers utilize floating tokens to build a global structural framework, while deeper layers allocate more capability toward capturing semantic content. Empirically, this distinctive attention pattern provides a mechanistic explanation for the strong in-context learning capabilities of MDMs, allowing them to double the performance compared to ARMs in knowledge-intensive tasks. All codes and datasets are available at https://github.com/NEUIR/Attention-Floating.",
"arxiv_id": "2601.07894",
"authors": [
"Xin Dai",
"Pengcheng Huang",
"Zhenghao Liu",
"Shuo Wang",
"Yukun Yan",
"Chaojun Xiao",
"Yu Gu",
"Ge Yu",
"Maosong Sun"
],
"categories": [
"cs.LG",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Revealing the Attention Floating Mechanism in Masked Diffusion Models",
"url": "https://arxiv.org/abs/2601.07894",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "3b781914-534f-4773-bbf8-1868b3de1c9b",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}