dorsal/arxiv
View SchemaOutcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning
| Authors | Ziheng Li, Liu Kang, Feng Xiao, Luxi Xing, Qingyi Si, Zhuoran Li, Weikang Gong, Deqing Yang, Yanghua Xiao, Hongcheng Guo |
|---|---|
| Categories | |
| ArXiv ID | 2601.07408vv1 |
| URL | https://arxiv.org/abs/2601.07408 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Group Relative Policy Optimization (GRPO) has emerged as a promising critic-free reinforcement learning paradigm for reasoning tasks. However, standard GRPO employs a coarse-grained credit assignment mechanism that propagates group-level rewards uniformly to to every token in a sequence, neglecting the varying contribution of individual reasoning steps. We address this limitation by introducing Outcome-grounded Advantage Reshaping (OAR), a fine-grained credit assignment mechanism that redistributes advantages based on how much each token influences the model's final answer. We instantiate OAR via two complementary strategies: (1) OAR-P, which estimates outcome sensitivity through counterfactual token perturbations, serving as a high-fidelity attribution signal; (2) OAR-G, which uses an input-gradient sensitivity proxy to approximate the influence signal with a single backward pass. These importance signals are integrated with a conservative Bi-Level advantage reshaping scheme that suppresses low-impact tokens and boosts pivotal ones while preserving the overall advantage mass. Empirical results on extensive mathematical reasoning benchmarks demonstrate that while OAR-P sets the performance upper bound, OAR-G achieves comparable gains with negligible computational overhead, both significantly outperforming a strong GRPO baseline, pushing the boundaries of critic-free LLM reasoning.
{
"annotation_id": "70bf40eb-fa76-480a-b8af-6b69a9785565",
"date_created": "2026-02-17T05:53:12.061000Z",
"date_modified": "2026-02-17T05:53:12.061000Z",
"file_hash": "eebca163bf61fea51840bc9e2969c31df8861fd01fca2190e9a5f15c5161ed4a",
"private": false,
"record": {
"abstract": "Group Relative Policy Optimization (GRPO) has emerged as a promising critic-free reinforcement learning paradigm for reasoning tasks. However, standard GRPO employs a coarse-grained credit assignment mechanism that propagates group-level rewards uniformly to to every token in a sequence, neglecting the varying contribution of individual reasoning steps. We address this limitation by introducing Outcome-grounded Advantage Reshaping (OAR), a fine-grained credit assignment mechanism that redistributes advantages based on how much each token influences the model\u0027s final answer. We instantiate OAR via two complementary strategies: (1) OAR-P, which estimates outcome sensitivity through counterfactual token perturbations, serving as a high-fidelity attribution signal; (2) OAR-G, which uses an input-gradient sensitivity proxy to approximate the influence signal with a single backward pass. These importance signals are integrated with a conservative Bi-Level advantage reshaping scheme that suppresses low-impact tokens and boosts pivotal ones while preserving the overall advantage mass. Empirical results on extensive mathematical reasoning benchmarks demonstrate that while OAR-P sets the performance upper bound, OAR-G achieves comparable gains with negligible computational overhead, both significantly outperforming a strong GRPO baseline, pushing the boundaries of critic-free LLM reasoning.",
"arxiv_id": "2601.07408",
"authors": [
"Ziheng Li",
"Liu Kang",
"Feng Xiao",
"Luxi Xing",
"Qingyi Si",
"Zhuoran Li",
"Weikang Gong",
"Deqing Yang",
"Yanghua Xiao",
"Hongcheng Guo"
],
"categories": [
"cs.CL",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Outcome-Grounded Advantage Reshaping for Fine-Grained Credit Assignment in Mathematical Reasoning",
"url": "https://arxiv.org/abs/2601.07408",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "2aafd005-7762-4928-ae8c-b24c4b4e203a",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}