dorsal/arxiv
View SchemaCollaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
| Authors | Zhiyuan Hu, Yunhai Hu, Juncheng Liu, Shuyue Stella Li, Yucheng Wang, Zhen Xu, See-Kiong Ng, Anh Tuan Luu, Xinxing Xu, Bryan Hooi, Cynthia Breazeal, Hae Won Park |
|---|---|
| Categories | |
| ArXiv ID | 2601.09667vv1 |
| URL | https://arxiv.org/abs/2601.09667 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Multi-agent systems have evolved into practical LLM-driven collaborators for many applications, gaining robustness from diversity and cross-checking. However, multi-agent RL (MARL) training is resource-intensive and unstable: co-adapting teammates induce non-stationarity, and rewards are often sparse and high-variance. Therefore, we introduce \textbf{Multi-Agent Test-Time Reinforcement Learning (MATTRL)}, a framework that injects structured textual experience into multi-agent deliberation at inference time. MATTRL forms a multi-expert team of specialists for multi-turn discussions, retrieves and integrates test-time experiences, and reaches consensus for final decision-making. We also study credit assignment for constructing a turn-level experience pool, then reinjecting it into the dialogue. Across challenging benchmarks in medicine, math, and education, MATTRL improves accuracy by an average of 3.67\% over a multi-agent baseline, and by 8.67\% over comparable single-agent baselines. Ablation studies examine different credit-assignment schemes and provide a detailed comparison of how they affect training outcomes. MATTRL offers a stable, effective and efficient path to distribution-shift-robust multi-agent reasoning without tuning.
{
"annotation_id": "c5dca186-93d9-4a6c-91a3-99303fc2da87",
"date_created": "2026-02-17T05:53:20.448000Z",
"date_modified": "2026-02-17T05:53:20.448000Z",
"file_hash": "f22be66fe25e4b8014704dca74f9e6c687f5afd4cf2e10ebcfef5f1c941671f5",
"private": false,
"record": {
"abstract": "Multi-agent systems have evolved into practical LLM-driven collaborators for many applications, gaining robustness from diversity and cross-checking. However, multi-agent RL (MARL) training is resource-intensive and unstable: co-adapting teammates induce non-stationarity, and rewards are often sparse and high-variance. Therefore, we introduce \\textbf{Multi-Agent Test-Time Reinforcement Learning (MATTRL)}, a framework that injects structured textual experience into multi-agent deliberation at inference time. MATTRL forms a multi-expert team of specialists for multi-turn discussions, retrieves and integrates test-time experiences, and reaches consensus for final decision-making. We also study credit assignment for constructing a turn-level experience pool, then reinjecting it into the dialogue. Across challenging benchmarks in medicine, math, and education, MATTRL improves accuracy by an average of 3.67\\% over a multi-agent baseline, and by 8.67\\% over comparable single-agent baselines. Ablation studies examine different credit-assignment schemes and provide a detailed comparison of how they affect training outcomes. MATTRL offers a stable, effective and efficient path to distribution-shift-robust multi-agent reasoning without tuning.",
"arxiv_id": "2601.09667",
"authors": [
"Zhiyuan Hu",
"Yunhai Hu",
"Juncheng Liu",
"Shuyue Stella Li",
"Yucheng Wang",
"Zhen Xu",
"See-Kiong Ng",
"Anh Tuan Luu",
"Xinxing Xu",
"Bryan Hooi",
"Cynthia Breazeal",
"Hae Won Park"
],
"categories": [
"cs.AI",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning",
"url": "https://arxiv.org/abs/2601.09667",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "153952d6-1440-434f-b5c7-8ffa77d1c71c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}