dorsal/arxiv
View SchemaSSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection
| Authors | Chenhao Fu, Han Fang, Xiuzheng Zheng, Wenbo Wei, Yonghua Li, Hao Sun, Xuelong Li |
|---|---|
| Categories | |
| ArXiv ID | 2601.09147vv1 |
| URL | https://arxiv.org/abs/2601.09147 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Zero-Shot Anomaly Detection (ZSAD) leverages Vision-Language Models (VLMs) to enable supervision-free industrial inspection. However, existing ZSAD paradigms are constrained by single visual backbones, which struggle to balance global semantic generalization with fine-grained structural discriminability. To bridge this gap, we propose Synergistic Semantic-Visual Prompting (SSVP), that efficiently fuses diverse visual encodings to elevate model's fine-grained perception. Specifically, SSVP introduces the Hierarchical Semantic-Visual Synergy (HSVS) mechanism, which deeply integrates DINOv3's multi-scale structural priors into the CLIP semantic space. Subsequently, the Vision-Conditioned Prompt Generator (VCPG) employs cross-modal attention to guide dynamic prompt generation, enabling linguistic queries to precisely anchor to specific anomaly patterns. Furthermore, to address the discrepancy between global scoring and local evidence, the Visual-Text Anomaly Mapper (VTAM) establishes a dual-gated calibration paradigm. Extensive evaluations on seven industrial benchmarks validate the robustness of our method; SSVP achieves state-of-the-art performance with 93.0\% Image-AUROC and 92.2\% Pixel-AUROC on MVTec-AD, significantly outperforming existing zero-shot approaches.
{
"annotation_id": "d316cd28-1396-4ceb-b427-639134b4ff77",
"date_created": "2026-02-17T05:53:20.070000Z",
"date_modified": "2026-02-17T05:53:20.070000Z",
"file_hash": "9041a15edf4e8bee75a21e7c3f4c129a43b64a7883ffb34b13547bb06eef725c",
"private": false,
"record": {
"abstract": "Zero-Shot Anomaly Detection (ZSAD) leverages Vision-Language Models (VLMs) to enable supervision-free industrial inspection. However, existing ZSAD paradigms are constrained by single visual backbones, which struggle to balance global semantic generalization with fine-grained structural discriminability. To bridge this gap, we propose Synergistic Semantic-Visual Prompting (SSVP), that efficiently fuses diverse visual encodings to elevate model\u0027s fine-grained perception. Specifically, SSVP introduces the Hierarchical Semantic-Visual Synergy (HSVS) mechanism, which deeply integrates DINOv3\u0027s multi-scale structural priors into the CLIP semantic space. Subsequently, the Vision-Conditioned Prompt Generator (VCPG) employs cross-modal attention to guide dynamic prompt generation, enabling linguistic queries to precisely anchor to specific anomaly patterns. Furthermore, to address the discrepancy between global scoring and local evidence, the Visual-Text Anomaly Mapper (VTAM) establishes a dual-gated calibration paradigm. Extensive evaluations on seven industrial benchmarks validate the robustness of our method; SSVP achieves state-of-the-art performance with 93.0\\% Image-AUROC and 92.2\\% Pixel-AUROC on MVTec-AD, significantly outperforming existing zero-shot approaches.",
"arxiv_id": "2601.09147",
"authors": [
"Chenhao Fu",
"Han Fang",
"Xiuzheng Zheng",
"Wenbo Wei",
"Yonghua Li",
"Hao Sun",
"Xuelong Li"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "SSVP: Synergistic Semantic-Visual Prompting for Industrial Zero-Shot Anomaly Detection",
"url": "https://arxiv.org/abs/2601.09147",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "db405b23-2a9d-4815-b564-d274bf59ad5c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}