dorsal/arxiv
View SchemaForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models
| Authors | Zhenhua Xu, Haobo Zhang, Zhebo Wang, Qichen Liu, Haitao Xu, Wenpeng Xing, Meng Han |
|---|---|
| Categories | |
| ArXiv ID | 2601.08189vv1 |
| URL | https://arxiv.org/abs/2601.08189 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Existing invasive (backdoor) fingerprints suffer from high-perplexity triggers that are easily filtered, fixed response patterns exposed by heuristic detectors, and spurious activations on benign inputs. We introduce \textsc{ForgetMark}, a stealthy fingerprinting framework that encodes provenance via targeted unlearning. It builds a compact, human-readable key--value set with an assistant model and predictive-entropy ranking, then trains lightweight LoRA adapters to suppress the original values on their keys while preserving general capabilities. Ownership is verified under black/gray-box access by aggregating likelihood and semantic evidence into a fingerprint success rate. By relying on probabilistic forgetting traces rather than fixed trigger--response patterns, \textsc{ForgetMark} avoids high-perplexity triggers, reduces detectability, and lowers false triggers. Across diverse architectures and settings, it achieves 100\% ownership verification on fingerprinted models while maintaining standard performance, surpasses backdoor baselines in stealthiness and robustness to model merging, and remains effective under moderate incremental fine-tuning. Our code and data are available at \href{https://github.com/Xuzhenhua55/ForgetMark}{https://github.com/Xuzhenhua55/ForgetMark}.
{
"annotation_id": "3abe288b-e221-49af-8fbf-6a107c027b5c",
"date_created": "2026-02-17T05:53:15.942000Z",
"date_modified": "2026-02-17T05:53:15.942000Z",
"file_hash": "8937f72d4c21931289488af6e540d65f1c46e8ade2b91892028726d47f019eda",
"private": false,
"record": {
"abstract": "Existing invasive (backdoor) fingerprints suffer from high-perplexity triggers that are easily filtered, fixed response patterns exposed by heuristic detectors, and spurious activations on benign inputs. We introduce \\textsc{ForgetMark}, a stealthy fingerprinting framework that encodes provenance via targeted unlearning. It builds a compact, human-readable key--value set with an assistant model and predictive-entropy ranking, then trains lightweight LoRA adapters to suppress the original values on their keys while preserving general capabilities. Ownership is verified under black/gray-box access by aggregating likelihood and semantic evidence into a fingerprint success rate. By relying on probabilistic forgetting traces rather than fixed trigger--response patterns, \\textsc{ForgetMark} avoids high-perplexity triggers, reduces detectability, and lowers false triggers. Across diverse architectures and settings, it achieves 100\\% ownership verification on fingerprinted models while maintaining standard performance, surpasses backdoor baselines in stealthiness and robustness to model merging, and remains effective under moderate incremental fine-tuning. Our code and data are available at \\href{https://github.com/Xuzhenhua55/ForgetMark}{https://github.com/Xuzhenhua55/ForgetMark}.",
"arxiv_id": "2601.08189",
"authors": [
"Zhenhua Xu",
"Haobo Zhang",
"Zhebo Wang",
"Qichen Liu",
"Haitao Xu",
"Wenpeng Xing",
"Meng Han"
],
"categories": [
"cs.CR",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "ForgetMark: Stealthy Fingerprint Embedding via Targeted Unlearning in Language Models",
"url": "https://arxiv.org/abs/2601.08189",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "159e4981-19d2-40ed-af87-1dd9c1419a57",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}