dorsal/arxiv
View SchemaCURVE: A Benchmark for Cultural and Multilingual Long Video Reasoning
| Authors | Darshan Singh, Arsha Nagrani, Kawshik Manikantan, Harman Singh, Dinesh Tewari, Tobias Weyand, Cordelia Schmid, Anelia Angelova, Shachi Dave |
|---|---|
| Categories | |
| ArXiv ID | 2601.10649vv1 |
| URL | https://arxiv.org/abs/2601.10649 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce CURVE (Cultural Understanding and Reasoning in Video Evaluation), a challenging benchmark for multicultural and multilingual video reasoning. CURVE comprises high-quality, entirely human-generated annotations from diverse, region-specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, CURVE provides complex questions, answers, and multi-step reasoning steps, all crafted in native languages. Making progress on CURVE requires a deeply situated understanding of visual cultural context. Furthermore, we leverage CURVE's reasoning traces to construct evidence-based graphs and propose a novel iterative strategy using these graphs to identify fine-grained errors in reasoning. Our evaluations reveal that SoTA Video-LLMs struggle significantly, performing substantially below human-level accuracy, with errors primarily stemming from the visual perception of cultural elements. CURVE will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\#minerva-cultural
{
"annotation_id": "9b4dcd2c-26d2-4362-b4d4-f40681ca3ad0",
"date_created": "2026-02-17T05:53:26.424000Z",
"date_modified": "2026-02-17T05:53:26.424000Z",
"file_hash": "ac705cb126239cabb18db652702186a75bb345a90e6a2ed445f0e1a36ae21b5d",
"private": false,
"record": {
"abstract": "Recent advancements in video models have shown tremendous progress, particularly in long video understanding. However, current benchmarks predominantly feature western-centric data and English as the dominant language, introducing significant biases in evaluation. To address this, we introduce CURVE (Cultural Understanding and Reasoning in Video Evaluation), a challenging benchmark for multicultural and multilingual video reasoning. CURVE comprises high-quality, entirely human-generated annotations from diverse, region-specific cultural videos across 18 global locales. Unlike prior work that relies on automatic translations, CURVE provides complex questions, answers, and multi-step reasoning steps, all crafted in native languages. Making progress on CURVE requires a deeply situated understanding of visual cultural context. Furthermore, we leverage CURVE\u0027s reasoning traces to construct evidence-based graphs and propose a novel iterative strategy using these graphs to identify fine-grained errors in reasoning. Our evaluations reveal that SoTA Video-LLMs struggle significantly, performing substantially below human-level accuracy, with errors primarily stemming from the visual perception of cultural elements. CURVE will be publicly available under https://github.com/google-deepmind/neptune?tab=readme-ov-file\\#minerva-cultural",
"arxiv_id": "2601.10649",
"authors": [
"Darshan Singh",
"Arsha Nagrani",
"Kawshik Manikantan",
"Harman Singh",
"Dinesh Tewari",
"Tobias Weyand",
"Cordelia Schmid",
"Anelia Angelova",
"Shachi Dave"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "CURVE: A Benchmark for Cultural and Multilingual Long Video Reasoning",
"url": "https://arxiv.org/abs/2601.10649",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "62744a92-41a2-4c2f-a1df-20b99ce07435",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}