dorsal/arxiv
View SchemaiReasoner: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models
| Authors | Meghana Sunil, Manikandarajan Venmathimaran, Muthu Subash Kavitha |
|---|---|
| Categories | |
| ArXiv ID | 2601.05877vv2 |
| URL | https://arxiv.org/abs/2601.05877 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Recent work shows that large multimodal models (LMMs) can self-improve from unlabeled data via self-play and intrinsic feedback. Yet existing self-evolving frameworks mainly reward final outcomes, leaving intermediate reasoning weakly constrained despite its importance for visually grounded decision making. We propose iReasoner, a self-evolving framework that improves an LMM's implicit reasoning by explicitly eliciting chain-of-thought (CoT) and rewarding its internal agreement. In a Proposer--Solver loop over unlabeled images, iReasoner augments outcome-level intrinsic rewards with a trajectory-aware signal defined over intermediate reasoning steps, providing learning signals that distinguish reasoning paths leading to the same answer without ground-truth labels or external judges. Starting from Qwen2.5-VL-7B, iReasoner yields up to $+2.1$ points across diverse multimodal reasoning benchmarks under fully unsupervised post-training. We hope this work serves as a starting point for reasoning-aware self-improvement in LMMs in purely unsupervised settings.
{
"annotation_id": "045eb7d7-b85a-4e0c-ae25-43882c2237dc",
"date_created": "2026-02-17T05:53:04.947000Z",
"date_modified": "2026-02-17T05:53:04.947000Z",
"file_hash": "40a6b211b565d5278c1c7c2f83fa166fa3be15321122956986e3d3a3f34d3493",
"private": false,
"record": {
"abstract": "Recent work shows that large multimodal models (LMMs) can self-improve from unlabeled data via self-play and intrinsic feedback. Yet existing self-evolving frameworks mainly reward final outcomes, leaving intermediate reasoning weakly constrained despite its importance for visually grounded decision making. We propose iReasoner, a self-evolving framework that improves an LMM\u0027s implicit reasoning by explicitly eliciting chain-of-thought (CoT) and rewarding its internal agreement. In a Proposer--Solver loop over unlabeled images, iReasoner augments outcome-level intrinsic rewards with a trajectory-aware signal defined over intermediate reasoning steps, providing learning signals that distinguish reasoning paths leading to the same answer without ground-truth labels or external judges. Starting from Qwen2.5-VL-7B, iReasoner yields up to $+2.1$ points across diverse multimodal reasoning benchmarks under fully unsupervised post-training. We hope this work serves as a starting point for reasoning-aware self-improvement in LMMs in purely unsupervised settings.",
"arxiv_id": "2601.05877",
"authors": [
"Meghana Sunil",
"Manikandarajan Venmathimaran",
"Muthu Subash Kavitha"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "iReasoner: Trajectory-Aware Intrinsic Reasoning Supervision for Self-Evolving Large Multimodal Models",
"url": "https://arxiv.org/abs/2601.05877",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "ec9744eb-5e90-427f-a180-c0669a78db40",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}