dorsal/arxiv
View SchemaSSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection
| Authors | Chenhao Fu, Han Fang, Xiuzheng Zheng, Wenbo Wei, Yonghua Li, Hao Sun, Xuelong Li |
|---|---|
| Categories | |
| ArXiv ID | 2601.09147vv2 |
| URL | https://arxiv.org/abs/2601.09147 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Zero-Shot Anomaly Detection (ZSAD) leverages Vision-Language Models (VLMs) to enable supervision-free industrial inspection. However, existing ZSAD paradigms are constrained by single visual backbones, which struggle to balance global semantic generalization with fine-grained structural discriminability. To bridge this gap, we propose Synergistic Semantic-Visual Prompting (SSVP), that efficiently fuses diverse visual encodings to elevate model's fine-grained perception. Specifically, SSVP introduces the Hierarchical Semantic-Visual Synergy (HSVS) mechanism, which deeply integrates DINOv3's multi-scale structural priors into the CLIP semantic space. Subsequently, the Vision-Conditioned Prompt Generator (VCPG) employs cross-modal attention to guide dynamic prompt generation, enabling linguistic queries to precisely anchor to specific anomaly patterns. Furthermore, to address the discrepancy between global scoring and local evidence, the Visual-Text Anomaly Mapper (VTAM) establishes a dual-gated calibration paradigm. Extensive evaluations on seven industrial benchmarks validate the robustness of our method; SSVP achieves state-of-the-art performance with 93.0\% Image-AUROC and 92.2\% Pixel-AUROC on MVTec-AD, significantly outperforming existing zero-shot approaches.
{
"annotation_id": "a421e0db-bf10-47ef-b616-a387171a3015",
"date_created": "2026-02-17T05:53:20.212000Z",
"date_modified": "2026-02-17T05:53:20.212000Z",
"file_hash": "4569a7dcdc4866ba0a9fc75ee43879f3f298725691fede5b115e959c4fe6ac10",
"private": false,
"record": {
"abstract": "Zero-Shot Anomaly Detection (ZSAD) leverages Vision-Language Models (VLMs) to enable supervision-free industrial inspection. However, existing ZSAD paradigms are constrained by single visual backbones, which struggle to balance global semantic generalization with fine-grained structural discriminability. To bridge this gap, we propose Synergistic Semantic-Visual Prompting (SSVP), that efficiently fuses diverse visual encodings to elevate model\u0027s fine-grained perception. Specifically, SSVP introduces the Hierarchical Semantic-Visual Synergy (HSVS) mechanism, which deeply integrates DINOv3\u0027s multi-scale structural priors into the CLIP semantic space. Subsequently, the Vision-Conditioned Prompt Generator (VCPG) employs cross-modal attention to guide dynamic prompt generation, enabling linguistic queries to precisely anchor to specific anomaly patterns. Furthermore, to address the discrepancy between global scoring and local evidence, the Visual-Text Anomaly Mapper (VTAM) establishes a dual-gated calibration paradigm. Extensive evaluations on seven industrial benchmarks validate the robustness of our method; SSVP achieves state-of-the-art performance with 93.0\\% Image-AUROC and 92.2\\% Pixel-AUROC on MVTec-AD, significantly outperforming existing zero-shot approaches.",
"arxiv_id": "2601.09147",
"authors": [
"Chenhao Fu",
"Han Fang",
"Xiuzheng Zheng",
"Wenbo Wei",
"Yonghua Li",
"Hao Sun",
"Xuelong Li"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection",
"url": "https://arxiv.org/abs/2601.09147",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d15073ee-380e-405c-ae5a-2d368fcc7492",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}