dorsal/arxiv
View SchemaCASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation
| Authors | Chaoyu Li, Deeparghya Dutta Barua, Fei Tao, Pooyan Fazli |
|---|---|
| Categories | |
| ArXiv ID | 2601.08010vv1 |
| URL | https://arxiv.org/abs/2601.08010 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasoning remains unstable. Repeated sampling over the same input often produces divergent reasoning trajectories and inconsistent final predictions. To address this, we introduce two complementary approaches inspired by test-time scaling: (1) CASHEW, an inference-time framework that stabilizes reasoning by iteratively aggregating multiple candidate trajectories into higher-quality reasoning traces, with explicit visual verification filtering hallucinated steps and grounding reasoning in visual evidence, and (2) CASHEW-RL, a learned variant that internalizes this aggregation behavior within a single model. CASHEW-RL is trained using Group Sequence Policy Optimization (GSPO) with a composite reward that encourages correct answers grounded in minimal yet sufficient visual evidence, while adaptively allocating reasoning effort based on task difficulty. This training objective enables robust self-aggregation at inference. Extensive experiments on 13 image understanding, video understanding, and video reasoning benchmarks show significant performance improvements, including gains of up to +23.6 percentage points on ScienceQA and +8.1 percentage points on EgoSchema.
{
"annotation_id": "b8238cf3-9a31-4e7c-8a3b-4099397e9fe4",
"date_created": "2026-02-17T05:53:15.040000Z",
"date_modified": "2026-02-17T05:53:15.040000Z",
"file_hash": "7a9f895eccb22121ebfe07959547a91ef021a997c71b5d1245256dc2fcdaee98",
"private": false,
"record": {
"abstract": "Vision-language models achieve strong performance across a wide range of multimodal understanding and reasoning tasks, yet their multi-step reasoning remains unstable. Repeated sampling over the same input often produces divergent reasoning trajectories and inconsistent final predictions. To address this, we introduce two complementary approaches inspired by test-time scaling: (1) CASHEW, an inference-time framework that stabilizes reasoning by iteratively aggregating multiple candidate trajectories into higher-quality reasoning traces, with explicit visual verification filtering hallucinated steps and grounding reasoning in visual evidence, and (2) CASHEW-RL, a learned variant that internalizes this aggregation behavior within a single model. CASHEW-RL is trained using Group Sequence Policy Optimization (GSPO) with a composite reward that encourages correct answers grounded in minimal yet sufficient visual evidence, while adaptively allocating reasoning effort based on task difficulty. This training objective enables robust self-aggregation at inference. Extensive experiments on 13 image understanding, video understanding, and video reasoning benchmarks show significant performance improvements, including gains of up to +23.6 percentage points on ScienceQA and +8.1 percentage points on EgoSchema.",
"arxiv_id": "2601.08010",
"authors": [
"Chaoyu Li",
"Deeparghya Dutta Barua",
"Fei Tao",
"Pooyan Fazli"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation",
"url": "https://arxiv.org/abs/2601.08010",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "f1db5c18-0ce5-4ea0-944b-1ac974bf65af",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}