dorsal/arxiv
View SchemaFast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning
| Authors | Chi-Pin Huang, Yunze Man, Zhiding Yu, Min-Hung Chen, Jan Kautz, Yu-Chiang Frank Wang, Fu-En Yang |
|---|---|
| Categories | |
| ArXiv ID | 2601.09708vv1 |
| URL | https://arxiv.org/abs/2601.09708 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Vision-Language-Action (VLA) tasks require reasoning over complex visual scenes and executing adaptive actions in dynamic environments. While recent studies on reasoning VLAs show that explicit chain-of-thought (CoT) can improve generalization, they suffer from high inference latency due to lengthy reasoning traces. We propose Fast-ThinkAct, an efficient reasoning framework that achieves compact yet performant planning through verbalizable latent reasoning. Fast-ThinkAct learns to reason efficiently with latent CoTs by distilling from a teacher, driven by a preference-guided objective to align manipulation trajectories that transfers both linguistic and visual planning capabilities for embodied control. This enables reasoning-enhanced policy learning that effectively connects compact reasoning to action execution. Extensive experiments across diverse embodied manipulation and reasoning benchmarks demonstrate that Fast-ThinkAct achieves strong performance with up to 89.3\% reduced inference latency over state-of-the-art reasoning VLAs, while maintaining effective long-horizon planning, few-shot adaptation, and failure recovery.
{
"annotation_id": "42772d33-d0c8-4bfe-9c6e-7b2470b442fa",
"date_created": "2026-02-17T05:53:20.177000Z",
"date_modified": "2026-02-17T05:53:20.177000Z",
"file_hash": "6ed51b092c1ef9e0b53c86df08304e1de8a19024ecf3938c98d1d614b2baae5f",
"private": false,
"record": {
"abstract": "Vision-Language-Action (VLA) tasks require reasoning over complex visual scenes and executing adaptive actions in dynamic environments. While recent studies on reasoning VLAs show that explicit chain-of-thought (CoT) can improve generalization, they suffer from high inference latency due to lengthy reasoning traces. We propose Fast-ThinkAct, an efficient reasoning framework that achieves compact yet performant planning through verbalizable latent reasoning. Fast-ThinkAct learns to reason efficiently with latent CoTs by distilling from a teacher, driven by a preference-guided objective to align manipulation trajectories that transfers both linguistic and visual planning capabilities for embodied control. This enables reasoning-enhanced policy learning that effectively connects compact reasoning to action execution. Extensive experiments across diverse embodied manipulation and reasoning benchmarks demonstrate that Fast-ThinkAct achieves strong performance with up to 89.3\\% reduced inference latency over state-of-the-art reasoning VLAs, while maintaining effective long-horizon planning, few-shot adaptation, and failure recovery.",
"arxiv_id": "2601.09708",
"authors": [
"Chi-Pin Huang",
"Yunze Man",
"Zhiding Yu",
"Min-Hung Chen",
"Jan Kautz",
"Yu-Chiang Frank Wang",
"Fu-En Yang"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.LG",
"cs.RO"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Fast-ThinkAct: Efficient Vision-Language-Action Reasoning via Verbalizable Latent Planning",
"url": "https://arxiv.org/abs/2601.09708",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "2e53381d-d1e3-4c78-a17d-28dca9c3faf0",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}