dorsal/arxiv
View SchemaOptimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge
| Authors | Sicheng Yang, Yukai Huang, Shitong Sun, Weitong Cai, Jiankang Deng, Jifei Song, Zhensong Zhang |
|---|---|
| Categories | |
| ArXiv ID | 2601.10228vv1 |
| URL | https://arxiv.org/abs/2601.10228 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD-EPIC VQA due to ambiguous queries/options, poor long-range temporal reasoning, and non-standardized outputs. We propose a framework integrating query/choice pre-processing, domain-specific Qwen2.5-VL fine-tuning, a novel Temporal Chain-of-Thought (T-CoT) prompting for multi-step reasoning, and robust post-processing. This system achieves 41.6% accuracy on HD-EPIC VQA, highlighting the need for holistic pipeline optimization in demanding video understanding. Our code, fine-tuned models are available at https://github.com/YoungSeng/Egocentric-Co-Pilot.
{
"annotation_id": "f89ef6b0-c19f-4085-8d6c-395a9209fcf7",
"date_created": "2026-02-17T05:53:23.615000Z",
"date_modified": "2026-02-17T05:53:23.615000Z",
"file_hash": "d4c5283ac60d168ec8cb82141613bbb3c6f35e539a28497663839204b432b999",
"private": false,
"record": {
"abstract": "Multimodal Large Language Models (MLLMs) struggle with complex video QA benchmarks like HD-EPIC VQA due to ambiguous queries/options, poor long-range temporal reasoning, and non-standardized outputs. We propose a framework integrating query/choice pre-processing, domain-specific Qwen2.5-VL fine-tuning, a novel Temporal Chain-of-Thought (T-CoT) prompting for multi-step reasoning, and robust post-processing. This system achieves 41.6% accuracy on HD-EPIC VQA, highlighting the need for holistic pipeline optimization in demanding video understanding. Our code, fine-tuned models are available at https://github.com/YoungSeng/Egocentric-Co-Pilot.",
"arxiv_id": "2601.10228",
"authors": [
"Sicheng Yang",
"Yukai Huang",
"Shitong Sun",
"Weitong Cai",
"Jiankang Deng",
"Jifei Song",
"Zhensong Zhang"
],
"categories": [
"cs.CV",
"cs.MM",
"eess.IV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Optimizing Multimodal LLMs for Egocentric Video Understanding: A Solution for the HD-EPIC VQA Challenge",
"url": "https://arxiv.org/abs/2601.10228",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "92c52b7d-ff53-41c9-a77e-a734d17fca41",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}