dorsal/arxiv
View SchemaSnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices
| Authors | Dongting Hu, Aarush Gupta, Magzhan Gabidolla, Arpit Sahni, Huseyin Coskun, Yanyu Li, Yerlan Idelbayev, Ahsan Mahmood, Aleksei Lebedev, Dishani Lahiri, Anujraaj Goyal, Ju Hu, Mingming Gong, Sergey Tulyakov, Anil Kag |
|---|---|
| Categories | |
| ArXiv ID | 2601.08303vv1 |
| URL | https://arxiv.org/abs/2601.08303 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to their high computational and memory costs. In this work, we present an efficient DiT framework tailored for mobile and edge devices that achieves transformer-level generation quality under strict resource constraints. Our design combines three key components. First, we propose a compact DiT architecture with an adaptive global-local sparse attention mechanism that balances global context modeling and local detail preservation. Second, we propose an elastic training framework that jointly optimizes sub-DiTs of varying capacities within a unified supernetwork, allowing a single model to dynamically adjust for efficient inference across different hardware. Finally, we develop Knowledge-Guided Distribution Matching Distillation, a step-distillation pipeline that integrates the DMD objective with knowledge transfer from few-step teacher models, producing high-fidelity and low-latency generation (e.g., 4-step) suitable for real-time on-device use. Together, these contributions enable scalable, efficient, and high-quality diffusion models for deployment on diverse hardware.
{
"annotation_id": "f4348d92-2da1-48e9-8d2f-9cc877772d3e",
"date_created": "2026-02-17T05:53:16.284000Z",
"date_modified": "2026-02-17T05:53:16.284000Z",
"file_hash": "d1e5d840fb7969d7579ca89cfd21fca262bb0826454b4b7cc3b92ec28d3412e5",
"private": false,
"record": {
"abstract": "Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to their high computational and memory costs. In this work, we present an efficient DiT framework tailored for mobile and edge devices that achieves transformer-level generation quality under strict resource constraints. Our design combines three key components. First, we propose a compact DiT architecture with an adaptive global-local sparse attention mechanism that balances global context modeling and local detail preservation. Second, we propose an elastic training framework that jointly optimizes sub-DiTs of varying capacities within a unified supernetwork, allowing a single model to dynamically adjust for efficient inference across different hardware. Finally, we develop Knowledge-Guided Distribution Matching Distillation, a step-distillation pipeline that integrates the DMD objective with knowledge transfer from few-step teacher models, producing high-fidelity and low-latency generation (e.g., 4-step) suitable for real-time on-device use. Together, these contributions enable scalable, efficient, and high-quality diffusion models for deployment on diverse hardware.",
"arxiv_id": "2601.08303",
"authors": [
"Dongting Hu",
"Aarush Gupta",
"Magzhan Gabidolla",
"Arpit Sahni",
"Huseyin Coskun",
"Yanyu Li",
"Yerlan Idelbayev",
"Ahsan Mahmood",
"Aleksei Lebedev",
"Dishani Lahiri",
"Anujraaj Goyal",
"Ju Hu",
"Mingming Gong",
"Sergey Tulyakov",
"Anil Kag"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices",
"url": "https://arxiv.org/abs/2601.08303",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6a1b5e06-1cb2-4b3b-a4dd-5eb81551cb43",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}