dorsal/arxiv
View SchemaCross-modal Proxy Evolving for OOD Detection with Vision-Language Models
| Authors | Hao Tang, Yu Liu, Shuanglin Yan, Fei Shen, Shengfeng He, Jing Qin |
|---|---|
| Categories | |
| ArXiv ID | 2601.08476vv1 |
| URL | https://arxiv.org/abs/2601.08476 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a proxy-aligned co-evolution mechanism to maintain two evolving proxy caches, which dynamically mines contextual textual negatives guided by test images and iteratively refines visual proxies, progressively realigning cross-modal similarities and enlarging local OOD margins. Finally, we dynamically re-weight the contributions of dual-modal proxies to obtain a calibrated OOD score that is robust to distribution shift. Extensive experiments on standard benchmarks demonstrate that CoEvo achieves state-of-the-art performance, improving AUROC by 1.33% and reducing FPR95 by 45.98% on ImageNet-1K compared to strong negative-label baselines.
{
"annotation_id": "6e6a3e4c-5634-4b13-8420-47813440b85d",
"date_created": "2026-02-17T05:53:15.844000Z",
"date_modified": "2026-02-17T05:53:15.844000Z",
"file_hash": "d41b44b0993712e38f01ff26f051d6d7c3b73c7e995f7bf61dbd51c7b714629e",
"private": false,
"record": {
"abstract": "Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in zero-shot OOD detection necessitates proxy signals that remain effective under distribution shift. Existing negative-label methods rely on a fixed set of textual proxies, which (i) sparsely sample the semantic space beyond in-distribution (ID) classes and (ii) remain static while only visual features drift, leading to cross-modal misalignment and unstable predictions. In this paper, we propose CoEvo, a training- and annotation-free test-time framework that performs bidirectional, sample-conditioned adaptation of both textual and visual proxies. Specifically, CoEvo introduces a proxy-aligned co-evolution mechanism to maintain two evolving proxy caches, which dynamically mines contextual textual negatives guided by test images and iteratively refines visual proxies, progressively realigning cross-modal similarities and enlarging local OOD margins. Finally, we dynamically re-weight the contributions of dual-modal proxies to obtain a calibrated OOD score that is robust to distribution shift. Extensive experiments on standard benchmarks demonstrate that CoEvo achieves state-of-the-art performance, improving AUROC by 1.33% and reducing FPR95 by 45.98% on ImageNet-1K compared to strong negative-label baselines.",
"arxiv_id": "2601.08476",
"authors": [
"Hao Tang",
"Yu Liu",
"Shuanglin Yan",
"Fei Shen",
"Shengfeng He",
"Jing Qin"
],
"categories": [
"cs.CV",
"cs.MM"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models",
"url": "https://arxiv.org/abs/2601.08476",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "2bb73f2f-2ced-4c15-bbc1-e29e0cd71b05",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}