dorsal/arxiv
View SchemaEnhancing Visual In-Context Learning by Multi-Faceted Fusion
| Authors | Wenwen Liao, Jianbo Yu, Yuansong Wang, Qingchao Jiang, Xiaofeng Yang |
|---|---|
| Categories | |
| ArXiv ID | 2601.10107vv1 |
| URL | https://arxiv.org/abs/2601.10107 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant "retrieve-then-prompt" approach typically relies on selecting the single best visual prompt, a practice that often discards valuable contextual information from other suitable candidates. While recent work has explored fusing the top-K prompts into a single, enhanced representation, this still simply collapses multiple rich signals into one, limiting the model's reasoning capability. We argue that a more multi-faceted, collaborative fusion is required to unlock the full potential of these diverse contexts. To address this limitation, we introduce a novel framework that moves beyond single-prompt fusion towards an multi-combination collaborative fusion. Instead of collapsing multiple prompts into one, our method generates three contextual representation branches, each formed by integrating information from different combinations of top-quality prompts. These complementary guidance signals are then fed into proposed MULTI-VQGAN architecture, which is designed to jointly interpret and utilize collaborative information from multiple sources. Extensive experiments on diverse tasks, including foreground segmentation, single-object detection, and image colorization, highlight its strong cross-task generalization, effective contextual fusion, and ability to produce more robust and accurate predictions than existing methods.
{
"annotation_id": "a206e918-63a8-42fd-96c3-e760799b1c8d",
"date_created": "2026-02-17T05:53:24.101000Z",
"date_modified": "2026-02-17T05:53:24.101000Z",
"file_hash": "e40071916a3527039042937c7260385cd99d06492b859c1995d3fea902d18f2a",
"private": false,
"record": {
"abstract": "Visual In-Context Learning (VICL) has emerged as a powerful paradigm, enabling models to perform novel visual tasks by learning from in-context examples. The dominant \"retrieve-then-prompt\" approach typically relies on selecting the single best visual prompt, a practice that often discards valuable contextual information from other suitable candidates. While recent work has explored fusing the top-K prompts into a single, enhanced representation, this still simply collapses multiple rich signals into one, limiting the model\u0027s reasoning capability. We argue that a more multi-faceted, collaborative fusion is required to unlock the full potential of these diverse contexts. To address this limitation, we introduce a novel framework that moves beyond single-prompt fusion towards an multi-combination collaborative fusion. Instead of collapsing multiple prompts into one, our method generates three contextual representation branches, each formed by integrating information from different combinations of top-quality prompts. These complementary guidance signals are then fed into proposed MULTI-VQGAN architecture, which is designed to jointly interpret and utilize collaborative information from multiple sources. Extensive experiments on diverse tasks, including foreground segmentation, single-object detection, and image colorization, highlight its strong cross-task generalization, effective contextual fusion, and ability to produce more robust and accurate predictions than existing methods.",
"arxiv_id": "2601.10107",
"authors": [
"Wenwen Liao",
"Jianbo Yu",
"Yuansong Wang",
"Qingchao Jiang",
"Xiaofeng Yang"
],
"categories": [
"cs.CV"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Enhancing Visual In-Context Learning by Multi-Faceted Fusion",
"url": "https://arxiv.org/abs/2601.10107",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "41f42328-0684-4346-a1f5-1cc54d0cdbfe",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}