dorsal/arxiv
View SchemaMMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting
| Authors | Kangda Wei, Ruihong Huang |
|---|---|
| Categories | |
| ArXiv ID | 2601.09085vv1 |
| URL | https://arxiv.org/abs/2601.09085 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
Group Relative Policy Optimization (GRPO) has become a standard approach for training mathematical reasoning models; however, its reliance on multiple completions per prompt makes training computationally expensive. Although recent work has reduced the number of training steps required to reach peak performance, the overall wall-clock training time often remains unchanged or even increases due to higher per-step cost. We propose MMR-GRPO, which integrates Maximal Marginal Relevance to reweigh rewards based on completion diversity. Our key insight is that semantically redundant completions contribute limited marginal learning signal; prioritizing diverse solutions yields more informative updates and accelerates convergence. Extensive evaluations across three model sizes (1.5B, 7B, 8B), three GRPO variants, and five mathematical reasoning benchmarks show that MMR-GRPO achieves comparable peak performance while requiring on average 47.9% fewer training steps and 70.2% less wall-clock time. These gains are consistent across models, methods, and benchmarks. We will release our code, trained models, and experimental protocols.
{
"annotation_id": "19c1ed9f-32ac-4199-ac56-da7d68d67539",
"date_created": "2026-02-17T05:53:20.223000Z",
"date_modified": "2026-02-17T05:53:20.223000Z",
"file_hash": "89c8b5e5cac23effe27714e4e1cd85c34512750699679f2c6ae513962ef94acd",
"private": false,
"record": {
"abstract": "Group Relative Policy Optimization (GRPO) has become a standard approach for training mathematical reasoning models; however, its reliance on multiple completions per prompt makes training computationally expensive. Although recent work has reduced the number of training steps required to reach peak performance, the overall wall-clock training time often remains unchanged or even increases due to higher per-step cost. We propose MMR-GRPO, which integrates Maximal Marginal Relevance to reweigh rewards based on completion diversity. Our key insight is that semantically redundant completions contribute limited marginal learning signal; prioritizing diverse solutions yields more informative updates and accelerates convergence. Extensive evaluations across three model sizes (1.5B, 7B, 8B), three GRPO variants, and five mathematical reasoning benchmarks show that MMR-GRPO achieves comparable peak performance while requiring on average 47.9% fewer training steps and 70.2% less wall-clock time. These gains are consistent across models, methods, and benchmarks. We will release our code, trained models, and experimental protocols.",
"arxiv_id": "2601.09085",
"authors": [
"Kangda Wei",
"Ruihong Huang"
],
"categories": [
"cs.LG",
"cs.AI",
"cs.CL",
"cs.IR"
],
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "MMR-GRPO: Accelerating GRPO-Style Training through Diversity-Aware Reward Reweighting",
"url": "https://arxiv.org/abs/2601.09085",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8f8e7cb9-0c5f-4727-a05f-4cac09ea23c3",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}