dorsal/arxiv
View SchemaVideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
| Authors | Longbin Ji, Xiaoxiong Liu, Junyuan Shang, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang |
|---|---|
| Categories | |
| ArXiv ID | 2601.05966vv2 |
| URL | https://arxiv.org/abs/2601.05966 |
| License | http://creativecommons.org/licenses/by-nc-sa/4.0/ |
Abstract
Recent advances in video generation have been dominated by diffusion and flow-matching models, which produce high-quality results but remain computationally intensive and difficult to scale. In this work, we introduce VideoAR, the first large-scale Visual Autoregressive (VAR) framework for video generation that combines multi-scale next-frame prediction with autoregressive modeling. VideoAR disentangles spatial and temporal dependencies by integrating intra-frame VAR modeling with causal next-frame prediction, supported by a 3D multi-scale tokenizer that efficiently encodes spatio-temporal dynamics. To improve long-term consistency, we propose Multi-scale Temporal RoPE, Cross-Frame Error Correction, and Random Frame Mask, which collectively mitigate error propagation and stabilize temporal coherence. Our multi-stage pretraining pipeline progressively aligns spatial and temporal learning across increasing resolutions and durations. Empirically, VideoAR achieves new state-of-the-art results among autoregressive models, improving FVD on UCF-101 from 99.5 to 88.6 while reducing inference steps by over 10x, and reaching a VBench score of 81.74-competitive with diffusion-based models an order of magnitude larger. These results demonstrate that VideoAR narrows the performance gap between autoregressive and diffusion paradigms, offering a scalable, efficient, and temporally consistent foundation for future video generation research.
{
"annotation_id": "8480c39b-30cf-4497-984e-29f90c197506",
"date_created": "2026-02-17T05:53:05.111000Z",
"date_modified": "2026-02-17T05:53:05.111000Z",
"file_hash": "531cc96a0d98285d1ff51119bc0507021317fc15e0e32f002b0f5d16adfe9554",
"private": false,
"record": {
"abstract": "Recent advances in video generation have been dominated by diffusion and flow-matching models, which produce high-quality results but remain computationally intensive and difficult to scale. In this work, we introduce VideoAR, the first large-scale Visual Autoregressive (VAR) framework for video generation that combines multi-scale next-frame prediction with autoregressive modeling. VideoAR disentangles spatial and temporal dependencies by integrating intra-frame VAR modeling with causal next-frame prediction, supported by a 3D multi-scale tokenizer that efficiently encodes spatio-temporal dynamics. To improve long-term consistency, we propose Multi-scale Temporal RoPE, Cross-Frame Error Correction, and Random Frame Mask, which collectively mitigate error propagation and stabilize temporal coherence. Our multi-stage pretraining pipeline progressively aligns spatial and temporal learning across increasing resolutions and durations. Empirically, VideoAR achieves new state-of-the-art results among autoregressive models, improving FVD on UCF-101 from 99.5 to 88.6 while reducing inference steps by over 10x, and reaching a VBench score of 81.74-competitive with diffusion-based models an order of magnitude larger. These results demonstrate that VideoAR narrows the performance gap between autoregressive and diffusion paradigms, offering a scalable, efficient, and temporally consistent foundation for future video generation research.",
"arxiv_id": "2601.05966",
"authors": [
"Longbin Ji",
"Xiaoxiong Liu",
"Junyuan Shang",
"Shuohuan Wang",
"Yu Sun",
"Hua Wu",
"Haifeng Wang"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by-nc-sa/4.0/",
"title": "VideoAR: Autoregressive Video Generation via Next-Frame \u0026 Scale Prediction",
"url": "https://arxiv.org/abs/2601.05966",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "7180a572-9793-4d21-be77-22e1f51a6493",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}