dorsal/arxiv
View SchemaThinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy
| Authors | Shujian Gao, Yuan Wang, Jiangtao Yan, Zuxuan Wu, Yu-Gang Jiang |
|---|---|
| Categories | |
| ArXiv ID | 2601.06801vv1 |
| URL | https://arxiv.org/abs/2601.06801 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced reasoning capabilities in Large Language Models. However, adapting RLVR to multimodal domains suffers from a critical \textit{perception-reasoning decoupling}. Existing paradigms, driven by text-centric outcome rewards, reasoning in language medium, inadvertently encourage models to bypass visual perception. We empirically validate this through blind experiments: state-of-the-art policies maintain or surprisingly improve performance even when visual inputs are entirely removed. This reveals that these models degenerate into \textit{blind reasoners}, exploiting linguistic priors to generate plausible answers instead of attending to visual evidence. In response, we propose \textbf{Thinking with Deltas}, a framework driven by a \textbf{Differential Visual Reasoning Policy (DVRP)}. DVRP introduces intrinsic supervision via visual triplets, comprising original, masked, and perturbed inputs. It optimizes the model to maximize reasoning divergence from masked inputs (enforcing \textit{visual sensitivity}) while minimizing divergence from perturbed inputs (ensuring \textit{visual robustness}). By aligning reasoning variations strictly with the \textit{Delta} of visual information, DVRP inherently bolsters visual understanding capabilities and significantly outperforms state-of-the-art methods on both general and medical benchmarks, without requiring external annotations or auxiliary tools.
{
"annotation_id": "4abac1cd-03f6-45c9-84d0-7c1a9230c9b3",
"date_created": "2026-02-17T05:53:08.728000Z",
"date_modified": "2026-02-17T05:53:08.728000Z",
"file_hash": "d5aa1d0f0263f41805329790ce2b7c839c76640beb3851ba4c8791ea9a91f92a",
"private": false,
"record": {
"abstract": "Reinforcement Learning with Verifiable Rewards (RLVR) has significantly advanced reasoning capabilities in Large Language Models. However, adapting RLVR to multimodal domains suffers from a critical \\textit{perception-reasoning decoupling}. Existing paradigms, driven by text-centric outcome rewards, reasoning in language medium, inadvertently encourage models to bypass visual perception. We empirically validate this through blind experiments: state-of-the-art policies maintain or surprisingly improve performance even when visual inputs are entirely removed. This reveals that these models degenerate into \\textit{blind reasoners}, exploiting linguistic priors to generate plausible answers instead of attending to visual evidence. In response, we propose \\textbf{Thinking with Deltas}, a framework driven by a \\textbf{Differential Visual Reasoning Policy (DVRP)}. DVRP introduces intrinsic supervision via visual triplets, comprising original, masked, and perturbed inputs. It optimizes the model to maximize reasoning divergence from masked inputs (enforcing \\textit{visual sensitivity}) while minimizing divergence from perturbed inputs (ensuring \\textit{visual robustness}). By aligning reasoning variations strictly with the \\textit{Delta} of visual information, DVRP inherently bolsters visual understanding capabilities and significantly outperforms state-of-the-art methods on both general and medical benchmarks, without requiring external annotations or auxiliary tools.",
"arxiv_id": "2601.06801",
"authors": [
"Shujian Gao",
"Yuan Wang",
"Jiangtao Yan",
"Zuxuan Wu",
"Yu-Gang Jiang"
],
"categories": [
"cs.AI",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Thinking with Deltas: Incentivizing Reinforcement Learning via Differential Visual Reasoning Policy",
"url": "https://arxiv.org/abs/2601.06801",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "f5fa4820-94d1-45b5-b4fc-749a6a1abc86",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}