dorsal/arxiv
View SchemaVideo Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding
| Authors | Yanxiang Huang, Guohua Gao, Zhaoyang Wei, Jianyuan Ni |
|---|---|
| Categories | |
| ArXiv ID | 2601.07761vv1 |
| URL | https://arxiv.org/abs/2601.07761 |
| Journal | ICME 2026 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches. To resolve this, we introduce the Chain of Evidence (CoE), a novel framework that architecturally decouples and co-optimizes perceptual grounding and reasoning efficiency. CoE incorporates two core innovations: (1) A lightweight Evidence Grounding Module (EGM) that acts as a query-guided filter, dynamically identifying and extracting a compact set of high-fidelity visual evidence; and (2) An Evidence-Anchoring Protocol optimized via Reinforcement Learning. Crucially, we design a composite reward mechanism that enforces process alignment, compelling the model to strictly reference identified temporal anchors during deduction, thereby mitigating hallucinations. To enable this, we construct CoE-Instruct, a large-scale dataset (164k samples) featuring a novel dual-annotation schema for separate perception and reasoning supervision. Extensive experiments on five benchmarks, including Video-MME, MVBench, and VSI-Bench, demonstrate that CoE-enhanced models establish a new state-of-the-art. They significantly outperform existing methods in accuracy, proving CoE to be a powerful and practical paradigm for reliable video understanding.
{
"annotation_id": "d200d84b-8234-42a5-97e3-bc8865bdb7ee",
"date_created": "2026-02-17T05:53:11.947000Z",
"date_modified": "2026-02-17T05:53:11.947000Z",
"file_hash": "7004f2c4b8047204e5bb53bda03b6f4df0da414e3fbe17c3a3b21d8940421fde",
"private": false,
"record": {
"abstract": "Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches. To resolve this, we introduce the Chain of Evidence (CoE), a novel framework that architecturally decouples and co-optimizes perceptual grounding and reasoning efficiency. CoE incorporates two core innovations: (1) A lightweight Evidence Grounding Module (EGM) that acts as a query-guided filter, dynamically identifying and extracting a compact set of high-fidelity visual evidence; and (2) An Evidence-Anchoring Protocol optimized via Reinforcement Learning. Crucially, we design a composite reward mechanism that enforces process alignment, compelling the model to strictly reference identified temporal anchors during deduction, thereby mitigating hallucinations. To enable this, we construct CoE-Instruct, a large-scale dataset (164k samples) featuring a novel dual-annotation schema for separate perception and reasoning supervision. Extensive experiments on five benchmarks, including Video-MME, MVBench, and VSI-Bench, demonstrate that CoE-enhanced models establish a new state-of-the-art. They significantly outperform existing methods in accuracy, proving CoE to be a powerful and practical paradigm for reliable video understanding.",
"arxiv_id": "2601.07761",
"authors": [
"Yanxiang Huang",
"Guohua Gao",
"Zhaoyang Wei",
"Jianyuan Ni"
],
"categories": [
"cs.CV"
],
"journal_ref": "ICME 2026",
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding",
"url": "https://arxiv.org/abs/2601.07761",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "ca83aaf1-6134-45dd-b55d-eb8e74c85c40",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}