dorsal/arxiv
View SchemaWhat Do LLM Agents Know About Their World? Task2Quiz: A Paradigm for Studying Environment Understanding
| Authors | Siyuan Liu, Hongbang Yuan, Xinze Li, Ziyue Zhu, Yixin Cao, Yu-Gang Jiang |
|---|---|
| Categories | |
| ArXiv ID | 2601.09503vv1 |
| URL | https://arxiv.org/abs/2601.09503 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large language model (LLM) agents have demonstrated remarkable capabilities in complex decision-making and tool-use tasks, yet their ability to generalize across varying environments remains a under-examined concern. Current evaluation paradigms predominantly rely on trajectory-based metrics that measure task success, while failing to assess whether agents possess a grounded, transferable model of the environment. To address this gap, we propose Task-to-Quiz (T2Q), a deterministic and automated evaluation paradigm designed to decouple task execution from world-state understanding. We instantiate this paradigm in T2QBench, a suite comprising 30 environments and 1,967 grounded QA pairs across multiple difficulty levels. Our extensive experiments reveal that task success is often a poor proxy for environment understanding, and that current memory machanism can not effectively help agents acquire a grounded model of the environment. These findings identify proactive exploration and fine-grained state representation as primary bottlenecks, offering a robust foundation for developing more generalizable autonomous agents.
{
"annotation_id": "4452ac78-65e5-4ad2-9f8f-b3d2c7e32a35",
"date_created": "2026-02-17T05:53:20.459000Z",
"date_modified": "2026-02-17T05:53:20.459000Z",
"file_hash": "43db8fa3776c42f76163c5c0490b56978783a3323bb3e5a0c328be09d4c07f52",
"private": false,
"record": {
"abstract": "Large language model (LLM) agents have demonstrated remarkable capabilities in complex decision-making and tool-use tasks, yet their ability to generalize across varying environments remains a under-examined concern. Current evaluation paradigms predominantly rely on trajectory-based metrics that measure task success, while failing to assess whether agents possess a grounded, transferable model of the environment. To address this gap, we propose Task-to-Quiz (T2Q), a deterministic and automated evaluation paradigm designed to decouple task execution from world-state understanding. We instantiate this paradigm in T2QBench, a suite comprising 30 environments and 1,967 grounded QA pairs across multiple difficulty levels. Our extensive experiments reveal that task success is often a poor proxy for environment understanding, and that current memory machanism can not effectively help agents acquire a grounded model of the environment. These findings identify proactive exploration and fine-grained state representation as primary bottlenecks, offering a robust foundation for developing more generalizable autonomous agents.",
"arxiv_id": "2601.09503",
"authors": [
"Siyuan Liu",
"Hongbang Yuan",
"Xinze Li",
"Ziyue Zhu",
"Yixin Cao",
"Yu-Gang Jiang"
],
"categories": [
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "What Do LLM Agents Know About Their World? Task2Quiz: A Paradigm for Studying Environment Understanding",
"url": "https://arxiv.org/abs/2601.09503",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "859689d3-aeeb-4d7a-9abc-08c8ac4d57a9",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}