dorsal/arxiv
View SchemaVULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding
| Authors | Haorui Yu, Ramon Ruiz-Dolz, Diji Yang, Hang He, Fengrui Zhang, Qiufeng Yi |
|---|---|
| Categories | |
| ArXiv ID | 2601.07986vv1 |
| URL | https://arxiv.org/abs/2601.07986 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
We introduce VULCA-Bench, a multicultural art-critique benchmark for evaluating Vision-Language Models' (VLMs) cultural understanding beyond surface-level visual perception. Existing VLM benchmarks predominantly measure L1-L2 capabilities (object recognition, scene description, and factual question answering) while under-evaluate higher-order cultural interpretation. VULCA-Bench contains 7,410 matched image-critique pairs spanning eight cultural traditions, with Chinese-English bilingual coverage. We operationalise cultural understanding using a five-layer framework (L1-L5, from Visual Perception to Philosophical Aesthetics), instantiated as 225 culture-specific dimensions and supported by expert-written bilingual critiques. Our pilot results indicate that higher-layer reasoning (L3-L5) is consistently more challenging than visual and technical analysis (L1-L2). The dataset, evaluation scripts, and annotation tools are available under CC BY 4.0 in the supplementary materials.
{
"annotation_id": "6f1b4343-4487-40e8-8e54-fc8385ea9e85",
"date_created": "2026-02-17T05:53:15.044000Z",
"date_modified": "2026-02-17T05:53:15.044000Z",
"file_hash": "f78f3870955e09c9a7c004eda57a9d66e46d85749d7203fbb34bf169b1f7f733",
"private": false,
"record": {
"abstract": "We introduce VULCA-Bench, a multicultural art-critique benchmark for evaluating Vision-Language Models\u0027 (VLMs) cultural understanding beyond surface-level visual perception. Existing VLM benchmarks predominantly measure L1-L2 capabilities (object recognition, scene description, and factual question answering) while under-evaluate higher-order cultural interpretation. VULCA-Bench contains 7,410 matched image-critique pairs spanning eight cultural traditions, with Chinese-English bilingual coverage. We operationalise cultural understanding using a five-layer framework (L1-L5, from Visual Perception to Philosophical Aesthetics), instantiated as 225 culture-specific dimensions and supported by expert-written bilingual critiques. Our pilot results indicate that higher-layer reasoning (L3-L5) is consistently more challenging than visual and technical analysis (L1-L2). The dataset, evaluation scripts, and annotation tools are available under CC BY 4.0 in the supplementary materials.",
"arxiv_id": "2601.07986",
"authors": [
"Haorui Yu",
"Ramon Ruiz-Dolz",
"Diji Yang",
"Hang He",
"Fengrui Zhang",
"Qiufeng Yi"
],
"categories": [
"cs.CL",
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "VULCA-Bench: A Multicultural Vision-Language Benchmark for Evaluating Cultural Understanding",
"url": "https://arxiv.org/abs/2601.07986",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "99f88aec-d6a5-451a-80e7-251cb47cc16d",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}