dorsal/arxiv
View SchemaDiscovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees
| Authors | Kun Li, Zenan Xu, Junan Li, Zengrui Jin, Jinghao Deng, Zexuan Qiu, Bo Zhou |
|---|---|
| Categories | |
| ArXiv ID | 2601.08274vv2 |
| URL | https://arxiv.org/abs/2601.08274 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Tool-Integrated Reasoning has emerged as a key paradigm to augment Large Language Models (LLMs) with computational capabilities, yet integrating tool-use into long Chain-of-Thought (long CoT) remains underexplored, largely due to the scarcity of training data and the challenge of integrating tool-use without compromising the model's intrinsic long-chain reasoning. In this paper, we introduce DART (Discovery And Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees), a reinforcement learning framework that enables spontaneous tool-use during long CoT reasoning without human annotation. DART operates by constructing dynamic rollout trees during training to discover valid tool-use opportunities, branching out at promising positions to explore diverse tool-integrated trajectories. Subsequently, a tree-based process advantage estimation identifies and credits specific sub-trajectories where tool invocation positively contributes to the solution, effectively reinforcing these beneficial behaviors. Extensive experiments on challenging benchmarks like AIME and GPQA-Diamond demonstrate that DART significantly outperforms existing methods, successfully harmonizing tool execution with long CoT reasoning.
{
"annotation_id": "ab799821-2daf-470b-bb86-e5d0ed44d167",
"date_created": "2026-02-17T05:53:15.435000Z",
"date_modified": "2026-02-17T05:53:15.435000Z",
"file_hash": "c3f6a17237e8dda322f6afae4ca3ca520c6419c203d37d8294c72f3d13d1683d",
"private": false,
"record": {
"abstract": "Tool-Integrated Reasoning has emerged as a key paradigm to augment Large Language Models (LLMs) with computational capabilities, yet integrating tool-use into long Chain-of-Thought (long CoT) remains underexplored, largely due to the scarcity of training data and the challenge of integrating tool-use without compromising the model\u0027s intrinsic long-chain reasoning. In this paper, we introduce DART (Discovery And Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees), a reinforcement learning framework that enables spontaneous tool-use during long CoT reasoning without human annotation. DART operates by constructing dynamic rollout trees during training to discover valid tool-use opportunities, branching out at promising positions to explore diverse tool-integrated trajectories. Subsequently, a tree-based process advantage estimation identifies and credits specific sub-trajectories where tool invocation positively contributes to the solution, effectively reinforcing these beneficial behaviors. Extensive experiments on challenging benchmarks like AIME and GPQA-Diamond demonstrate that DART significantly outperforms existing methods, successfully harmonizing tool execution with long CoT reasoning.",
"arxiv_id": "2601.08274",
"authors": [
"Kun Li",
"Zenan Xu",
"Junan Li",
"Zengrui Jin",
"Jinghao Deng",
"Zexuan Qiu",
"Bo Zhou"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Discovery and Reinforcement of Tool-Integrated Reasoning Chains via Rollout Trees",
"url": "https://arxiv.org/abs/2601.08274",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8f9baf03-f27a-4fda-a012-9e5e3a31aa08",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}