dorsal/arxiv
View SchemaToolGym: an Open-world Tool-using Environment for Scalable Agent Testing and Data Curation
| Authors | Ziqiao Xi, Shuang Liang, Qi Liu, Jiaqing Zhang, Letian Peng, Fang Nan, Meshal Nayim, Tianhui Zhang, Rishika Mundada, Lianhui Qin, Biwei Huang, Kun Zhou |
|---|---|
| Categories | |
| ArXiv ID | 2601.06328vv1 |
| URL | https://arxiv.org/abs/2601.06328 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Tool-using LLM agents still struggle in open-world settings with large tool pools, long-horizon objectives, wild constraints, and unreliable tool states. For scalable and realistic training and testing, we introduce an open-world tool-using environment, built on 5,571 format unified tools across 204 commonly used apps. It includes a task creation engine that synthesizes long-horizon, multi-tool workflows with wild constraints, and a state controller that injects interruptions and failures to stress-test robustness. On top of this environment, we develop a tool select-then-execute agent framework with a planner-actor decomposition to separate deliberate reasoning and self-correction from step-wise execution. Comprehensive evaluation of state-of-the-art LLMs reveals the misalignment between tool planning and execution abilities, the constraint following weakness of existing LLMs, and DeepSeek-v3.2's strongest robustness. Finally, we collect 1,170 trajectories from our environment to fine-tune LLMs, achieving superior performance to baselines using 119k samples, indicating the environment's value as both a realistic benchmark and a data engine for tool-using agents. Our code and data will be publicly released.
{
"annotation_id": "b234fc9c-9318-460d-9019-14df57379a18",
"date_created": "2026-02-17T05:53:08.558000Z",
"date_modified": "2026-02-17T05:53:08.558000Z",
"file_hash": "ffa3a2543103727774ea8c0ace12ecf6269af9d3ad569e028ca2bea3e623ae6b",
"private": false,
"record": {
"abstract": "Tool-using LLM agents still struggle in open-world settings with large tool pools, long-horizon objectives, wild constraints, and unreliable tool states. For scalable and realistic training and testing, we introduce an open-world tool-using environment, built on 5,571 format unified tools across 204 commonly used apps. It includes a task creation engine that synthesizes long-horizon, multi-tool workflows with wild constraints, and a state controller that injects interruptions and failures to stress-test robustness. On top of this environment, we develop a tool select-then-execute agent framework with a planner-actor decomposition to separate deliberate reasoning and self-correction from step-wise execution. Comprehensive evaluation of state-of-the-art LLMs reveals the misalignment between tool planning and execution abilities, the constraint following weakness of existing LLMs, and DeepSeek-v3.2\u0027s strongest robustness. Finally, we collect 1,170 trajectories from our environment to fine-tune LLMs, achieving superior performance to baselines using 119k samples, indicating the environment\u0027s value as both a realistic benchmark and a data engine for tool-using agents. Our code and data will be publicly released.",
"arxiv_id": "2601.06328",
"authors": [
"Ziqiao Xi",
"Shuang Liang",
"Qi Liu",
"Jiaqing Zhang",
"Letian Peng",
"Fang Nan",
"Meshal Nayim",
"Tianhui Zhang",
"Rishika Mundada",
"Lianhui Qin",
"Biwei Huang",
"Kun Zhou"
],
"categories": [
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "ToolGym: an Open-world Tool-using Environment for Scalable Agent Testing and Data Curation",
"url": "https://arxiv.org/abs/2601.06328",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "06e67c7c-0139-490b-b85e-93e92553df54",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}