dorsal/arxiv
View SchemaTransition Matching Distillation for Fast Video Generation
| Authors | Weili Nie, Julius Berner, Nanye Ma, Chao Liu, Saining Xie, Arash Vahdat |
|---|---|
| Categories | |
| ArXiv ID | 2601.09881vv1 |
| URL | https://arxiv.org/abs/2601.09881 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process. In this work, we present Transition Matching Distillation (TMD), a novel framework for distilling video diffusion models into efficient few-step generators. The central idea of TMD is to match the multi-step denoising trajectory of a diffusion model with a few-step probability transition process, where each transition is modeled as a lightweight conditional flow. To enable efficient distillation, we decompose the original diffusion backbone into two components: (1) a main backbone, comprising the majority of early layers, that extracts semantic representations at each outer transition step; and (2) a flow head, consisting of the last few layers, that leverages these representations to perform multiple inner flow updates. Given a pretrained video diffusion model, we first introduce a flow head to the model, and adapt it into a conditional flow map. We then apply distribution matching distillation to the student model with flow head rollout in each transition step. Extensive experiments on distilling Wan2.1 1.3B and 14B text-to-video models demonstrate that TMD provides a flexible and strong trade-off between generation speed and visual quality. In particular, TMD outperforms existing distilled models under comparable inference costs in terms of visual fidelity and prompt adherence. Project page: https://research.nvidia.com/labs/genair/tmd
{
"annotation_id": "699e2401-f0b1-4d95-b902-1922c66f0668",
"date_created": "2026-02-17T05:53:24.322000Z",
"date_modified": "2026-02-17T05:53:24.322000Z",
"file_hash": "dfdc81bf259f64401d32ad3f697dc4b723413b25545865c9f5a12e27bf3e5bef",
"private": false,
"record": {
"abstract": "Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process. In this work, we present Transition Matching Distillation (TMD), a novel framework for distilling video diffusion models into efficient few-step generators. The central idea of TMD is to match the multi-step denoising trajectory of a diffusion model with a few-step probability transition process, where each transition is modeled as a lightweight conditional flow. To enable efficient distillation, we decompose the original diffusion backbone into two components: (1) a main backbone, comprising the majority of early layers, that extracts semantic representations at each outer transition step; and (2) a flow head, consisting of the last few layers, that leverages these representations to perform multiple inner flow updates. Given a pretrained video diffusion model, we first introduce a flow head to the model, and adapt it into a conditional flow map. We then apply distribution matching distillation to the student model with flow head rollout in each transition step. Extensive experiments on distilling Wan2.1 1.3B and 14B text-to-video models demonstrate that TMD provides a flexible and strong trade-off between generation speed and visual quality. In particular, TMD outperforms existing distilled models under comparable inference costs in terms of visual fidelity and prompt adherence. Project page: https://research.nvidia.com/labs/genair/tmd",
"arxiv_id": "2601.09881",
"authors": [
"Weili Nie",
"Julius Berner",
"Nanye Ma",
"Chao Liu",
"Saining Xie",
"Arash Vahdat"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Transition Matching Distillation for Fast Video Generation",
"url": "https://arxiv.org/abs/2601.09881",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "31d762bf-17f3-49b6-aa16-2a95f9b3aa80",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}