dorsal/arxiv
View SchemaV-Zero: Self-Improving Multimodal Reasoning with Zero Annotation
| Authors | Han Wang, Yi Yang, Jingyuan Hu, Minfeng Zhu, Wei Chen |
|---|---|
| Categories | |
| ArXiv ID | 2601.10094vv1 |
| URL | https://arxiv.org/abs/2601.10094 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision-language models (VLMs). However, state-of-the-art approaches rely heavily on large-scale human-annotated datasets, which are costly and time-consuming to acquire. To overcome this limitation, we introduce V-Zero, a general post-training framework that facilitates self-improvement using exclusively unlabeled images. V-Zero establishes a co-evolutionary loop by instantiating two distinct roles: a Questioner and a Solver. The Questioner learns to synthesize high-quality, challenging questions by leveraging a dual-track reasoning reward that contrasts intuitive guesses with reasoned results. The Solver is optimized using pseudo-labels derived from majority voting over its own sampled responses. Both roles are trained iteratively via Group Relative Policy Optimization (GRPO), driving a cycle of mutual enhancement. Remarkably, without a single human annotation, V-Zero achieves consistent performance gains on Qwen2.5-VL-7B-Instruct, improving visual mathematical reasoning by +1.7 and general vision-centric by +2.6, demonstrating the potential of self-improvement in multimodal systems. Code is available at https://github.com/SatonoDia/V-Zero
{
"annotation_id": "370e464c-1aad-46b3-a16c-95371c012df6",
"date_created": "2026-02-17T05:53:24.214000Z",
"date_modified": "2026-02-17T05:53:24.214000Z",
"file_hash": "2838df2857b5236e0d88a0becccf2d37c7c59906d4889b2362cde8788d8c6e69",
"private": false,
"record": {
"abstract": "Recent advances in multimodal learning have significantly enhanced the reasoning capabilities of vision-language models (VLMs). However, state-of-the-art approaches rely heavily on large-scale human-annotated datasets, which are costly and time-consuming to acquire. To overcome this limitation, we introduce V-Zero, a general post-training framework that facilitates self-improvement using exclusively unlabeled images. V-Zero establishes a co-evolutionary loop by instantiating two distinct roles: a Questioner and a Solver. The Questioner learns to synthesize high-quality, challenging questions by leveraging a dual-track reasoning reward that contrasts intuitive guesses with reasoned results. The Solver is optimized using pseudo-labels derived from majority voting over its own sampled responses. Both roles are trained iteratively via Group Relative Policy Optimization (GRPO), driving a cycle of mutual enhancement. Remarkably, without a single human annotation, V-Zero achieves consistent performance gains on Qwen2.5-VL-7B-Instruct, improving visual mathematical reasoning by +1.7 and general vision-centric by +2.6, demonstrating the potential of self-improvement in multimodal systems. Code is available at https://github.com/SatonoDia/V-Zero",
"arxiv_id": "2601.10094",
"authors": [
"Han Wang",
"Yi Yang",
"Jingyuan Hu",
"Minfeng Zhu",
"Wei Chen"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "V-Zero: Self-Improving Multimodal Reasoning with Zero Annotation",
"url": "https://arxiv.org/abs/2601.10094",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "3ce91960-22c2-4a7a-902c-5d702888409a",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}