dorsal/arxiv
View SchemaFigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures
| Authors | Jifeng Song, Arun Das, Pan Wang, Hui Ji, Kun Zhao, Yufei Huang |
|---|---|
| Categories | |
| ArXiv ID | 2601.08026vv2 |
| URL | https://arxiv.org/abs/2601.08026 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Scientific compound figures combine multiple labeled panels into a single image, but captions in real pipelines are often missing or only provide figure-level summaries, making panel-level understanding difficult. In this paper, we propose FigEx2, visual-conditioned framework that localizes panels and generates panel-wise captions directly from the compound figure. To mitigate the impact of diverse phrasing in open-ended captioning, we introduce a noise-aware gated fusion module that adaptively filters token-level features to stabilize the detection query space. Furthermore, we employ a staged optimization strategy combining supervised learning with reinforcement learning (RL), utilizing CLIP-based alignment and BERTScore-based semantic rewards to enforce strict multimodal consistency. To support high-quality supervision, we curate BioSci-Fig-Cap, a refined benchmark for panel-level grounding, alongside cross-disciplinary test suites in physics and chemistry. Experimental results demonstrate that FigEx2 achieves a superior 0.726 mAP@0.5:0.95 for detection and significantly outperforms Qwen3-VL-8B by 0.51 in METEOR and 0.24 in BERTScore. Notably, FigEx2 exhibits remarkable zero-shot transferability to out-of-distribution scientific domains without any fine-tuning.
{
"annotation_id": "49cc315a-1959-4949-b24f-8229c43e8d77",
"date_created": "2026-02-17T05:53:16.281000Z",
"date_modified": "2026-02-17T05:53:16.281000Z",
"file_hash": "bee72b36b80483381d2ad26dd4d626764f0b8357823df5b1e181c2837d3f46bd",
"private": false,
"record": {
"abstract": "Scientific compound figures combine multiple labeled panels into a single image, but captions in real pipelines are often missing or only provide figure-level summaries, making panel-level understanding difficult. In this paper, we propose FigEx2, visual-conditioned framework that localizes panels and generates panel-wise captions directly from the compound figure. To mitigate the impact of diverse phrasing in open-ended captioning, we introduce a noise-aware gated fusion module that adaptively filters token-level features to stabilize the detection query space. Furthermore, we employ a staged optimization strategy combining supervised learning with reinforcement learning (RL), utilizing CLIP-based alignment and BERTScore-based semantic rewards to enforce strict multimodal consistency. To support high-quality supervision, we curate BioSci-Fig-Cap, a refined benchmark for panel-level grounding, alongside cross-disciplinary test suites in physics and chemistry. Experimental results demonstrate that FigEx2 achieves a superior 0.726 mAP@0.5:0.95 for detection and significantly outperforms Qwen3-VL-8B by 0.51 in METEOR and 0.24 in BERTScore. Notably, FigEx2 exhibits remarkable zero-shot transferability to out-of-distribution scientific domains without any fine-tuning.",
"arxiv_id": "2601.08026",
"authors": [
"Jifeng Song",
"Arun Das",
"Pan Wang",
"Hui Ji",
"Kun Zhao",
"Yufei Huang"
],
"categories": [
"cs.CV",
"cs.AI",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "FigEx2: Visual-Conditioned Panel Detection and Captioning for Scientific Compound Figures",
"url": "https://arxiv.org/abs/2601.08026",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8f7a7856-c07f-4b73-8a12-964c3f80e4c0",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}