dorsal/arxiv
View SchemaImproving Video Question Answering through query-based frame selection
| Authors | Himanshu Patil, Geo Jolly, Ramana Raja Buddala, Ganesh Ramakrishnan, Rohit Saluja |
|---|---|
| Categories | |
| ArXiv ID | 2601.07459vv1 |
| URL | https://arxiv.org/abs/2601.07459 |
| DOI | 10.1145/3774521.3774607 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Video Question Answering (VideoQA) models enhance understanding and interaction with audiovisual content, making it more accessible, searchable, and useful for a wide range of fields such as education, surveillance, entertainment, and content creation. Due to heavy compute requirements, most large visual language models (VLMs) for VideoQA rely on a fixed number of frames by uniformly sampling the video. However, this process does not pick important frames or capture the context of the video. We present a novel query-based selection of frames relevant to the questions based on the submodular mutual Information (SMI) functions. By replacing uniform frame sampling with query-based selection, our method ensures that the chosen frames provide complementary and essential visual information for accurate VideoQA. We evaluate our approach on the MVBench dataset, which spans a diverse set of multi-action video tasks. VideoQA accuracy on this dataset was assessed using two VLMs, namely Video-LLaVA and LLaVA-NeXT, both of which originally employed uniform frame sampling. Experiments were conducted using both uniform and query-based sampling strategies. An accuracy improvement of up to \textbf{4\%} was observed when using query-based frame selection over uniform sampling. Qualitative analysis further highlights that query-based selection, using SMI functions, consistently picks frames better aligned with the question. We opine that such query-based frame selection can enhance accuracy in a wide range of tasks that rely on only a subset of video frames.
{
"annotation_id": "962baeec-5373-4166-ae25-ca42fcad7a51",
"date_created": "2026-02-17T05:53:11.557000Z",
"date_modified": "2026-02-17T05:53:11.557000Z",
"file_hash": "a5100c9be56516264965baa91f311cd067528f7c58e5ffd02b3765cd4c50f95b",
"private": false,
"record": {
"abstract": "Video Question Answering (VideoQA) models enhance understanding and interaction with audiovisual content, making it more accessible, searchable, and useful for a wide range of fields such as education, surveillance, entertainment, and content creation. Due to heavy compute requirements, most large visual language models (VLMs) for VideoQA rely on a fixed number of frames by uniformly sampling the video. However, this process does not pick important frames or capture the context of the video. We present a novel query-based selection of frames relevant to the questions based on the submodular mutual Information (SMI) functions. By replacing uniform frame sampling with query-based selection, our method ensures that the chosen frames provide complementary and essential visual information for accurate VideoQA. We evaluate our approach on the MVBench dataset, which spans a diverse set of multi-action video tasks. VideoQA accuracy on this dataset was assessed using two VLMs, namely Video-LLaVA and LLaVA-NeXT, both of which originally employed uniform frame sampling. Experiments were conducted using both uniform and query-based sampling strategies. An accuracy improvement of up to \\textbf{4\\%} was observed when using query-based frame selection over uniform sampling. Qualitative analysis further highlights that query-based selection, using SMI functions, consistently picks frames better aligned with the question. We opine that such query-based frame selection can enhance accuracy in a wide range of tasks that rely on only a subset of video frames.",
"arxiv_id": "2601.07459",
"authors": [
"Himanshu Patil",
"Geo Jolly",
"Ramana Raja Buddala",
"Ganesh Ramakrishnan",
"Rohit Saluja"
],
"categories": [
"cs.CV",
"cs.LG"
],
"doi": "10.1145/3774521.3774607",
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Improving Video Question Answering through query-based frame selection",
"url": "https://arxiv.org/abs/2601.07459",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e034a187-728b-4eec-b694-069a81dc97ef",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}