dorsal/arxiv
View SchemaDNF: Dual-Layer Nested Fingerprinting for Large Language Model Intellectual Property Protection
| Authors | Zhenhua Xu, Yiran Zhao, Mengting Zhong, Dezhang Kong, Changting Lin, Tong Qiao, Meng Han |
|---|---|
| Categories | |
| ArXiv ID | 2601.08223vv2 |
| URL | https://arxiv.org/abs/2601.08223 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The rapid growth of large language models raises pressing concerns about intellectual property protection under black-box deployment. Existing backdoor-based fingerprints either rely on rare tokens -- leading to high-perplexity inputs susceptible to filtering -- or use fixed trigger-response mappings that are brittle to leakage and post-hoc adaptation. We propose \textsc{Dual-Layer Nested Fingerprinting} (DNF), a black-box method that embeds a hierarchical backdoor by coupling domain-specific stylistic cues with implicit semantic triggers. Across Mistral-7B, LLaMA-3-8B-Instruct, and Falcon3-7B-Instruct, DNF achieves perfect fingerprint activation while preserving downstream utility. Compared with existing methods, it uses lower-perplexity triggers, remains undetectable under fingerprint detection attacks, and is relatively robust to incremental fine-tuning and model merging. These results position DNF as a practical, stealthy, and resilient solution for LLM ownership verification and intellectual property protection.
{
"annotation_id": "1e0c871d-63a7-4119-a952-033d6023d3f3",
"date_created": "2026-02-17T05:53:16.091000Z",
"date_modified": "2026-02-17T05:53:16.091000Z",
"file_hash": "bb3c87e7ddc01f1a4471aaa83072693ed0d75e548d5f6cae4a5553e8a738b4fd",
"private": false,
"record": {
"abstract": "The rapid growth of large language models raises pressing concerns about intellectual property protection under black-box deployment. Existing backdoor-based fingerprints either rely on rare tokens -- leading to high-perplexity inputs susceptible to filtering -- or use fixed trigger-response mappings that are brittle to leakage and post-hoc adaptation. We propose \\textsc{Dual-Layer Nested Fingerprinting} (DNF), a black-box method that embeds a hierarchical backdoor by coupling domain-specific stylistic cues with implicit semantic triggers. Across Mistral-7B, LLaMA-3-8B-Instruct, and Falcon3-7B-Instruct, DNF achieves perfect fingerprint activation while preserving downstream utility. Compared with existing methods, it uses lower-perplexity triggers, remains undetectable under fingerprint detection attacks, and is relatively robust to incremental fine-tuning and model merging. These results position DNF as a practical, stealthy, and resilient solution for LLM ownership verification and intellectual property protection.",
"arxiv_id": "2601.08223",
"authors": [
"Zhenhua Xu",
"Yiran Zhao",
"Mengting Zhong",
"Dezhang Kong",
"Changting Lin",
"Tong Qiao",
"Meng Han"
],
"categories": [
"cs.CR",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "DNF: Dual-Layer Nested Fingerprinting for Large Language Model Intellectual Property Protection",
"url": "https://arxiv.org/abs/2601.08223",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "a7248966-0b4c-46ca-b1fa-c57c9641c98f",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}