dorsal/arxiv
View SchemaAgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
| Authors | Xuannan Liu, Xiao Yang, Zekun Li, Peipei Li, Ran He |
|---|---|
| Categories | |
| ArXiv ID | 2601.06818vv1 |
| URL | https://arxiv.org/abs/2601.06818 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
As LLM-based agents operate over sequential multi-step reasoning, hallucinations arising at intermediate steps risk propagating along the trajectory, thus degrading overall reliability. Unlike hallucination detection in single-turn responses, diagnosing hallucinations in multi-step workflows requires identifying which step causes the initial divergence. To fill this gap, we propose a new research task, automated hallucination attribution of LLM-based agents, aiming to identify the step responsible for the hallucination and explain why. To support this task, we introduce AgentHallu, a comprehensive benchmark with: (1) 693 high-quality trajectories spanning 7 agent frameworks and 5 domains, (2) a hallucination taxonomy organized into 5 categories (Planning, Retrieval, Reasoning, Human-Interaction, and Tool-Use) and 14 sub-categories, and (3) multi-level annotations curated by humans, covering binary labels, hallucination-responsible steps, and causal explanations. We evaluate 13 leading models, and results show the task is challenging even for top-tier models (like GPT-5, Gemini-2.5-Pro). The best-performing model achieves only 41.1\% step localization accuracy, where tool-use hallucinations are the most challenging at just 11.6\%. We believe AgentHallu will catalyze future research into developing robust, transparent, and reliable agentic systems.
{
"annotation_id": "356d6f65-ae08-4759-891f-d6b73a29e0df",
"date_created": "2026-02-17T05:53:08.864000Z",
"date_modified": "2026-02-17T05:53:08.864000Z",
"file_hash": "811c95dbc1bbf59d5be523cecba77d702a3409e66fe5501937a0c0a44506c299",
"private": false,
"record": {
"abstract": "As LLM-based agents operate over sequential multi-step reasoning, hallucinations arising at intermediate steps risk propagating along the trajectory, thus degrading overall reliability. Unlike hallucination detection in single-turn responses, diagnosing hallucinations in multi-step workflows requires identifying which step causes the initial divergence. To fill this gap, we propose a new research task, automated hallucination attribution of LLM-based agents, aiming to identify the step responsible for the hallucination and explain why. To support this task, we introduce AgentHallu, a comprehensive benchmark with: (1) 693 high-quality trajectories spanning 7 agent frameworks and 5 domains, (2) a hallucination taxonomy organized into 5 categories (Planning, Retrieval, Reasoning, Human-Interaction, and Tool-Use) and 14 sub-categories, and (3) multi-level annotations curated by humans, covering binary labels, hallucination-responsible steps, and causal explanations. We evaluate 13 leading models, and results show the task is challenging even for top-tier models (like GPT-5, Gemini-2.5-Pro). The best-performing model achieves only 41.1\\% step localization accuracy, where tool-use hallucinations are the most challenging at just 11.6\\%. We believe AgentHallu will catalyze future research into developing robust, transparent, and reliable agentic systems.",
"arxiv_id": "2601.06818",
"authors": [
"Xuannan Liu",
"Xiao Yang",
"Zekun Li",
"Peipei Li",
"Ran He"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents",
"url": "https://arxiv.org/abs/2601.06818",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e496f03a-7386-44be-a559-20ee9b43dfa3",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}