dorsal/arxiv
View SchemaPDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection
| Authors | Jinhan Liu, Yibo Yang, Ruiying Lu, Piotr Piekos, Yimeng Chen, Peng Wang, Dandan Guo |
|---|---|
| Categories | |
| ArXiv ID | 2601.06827vv1 |
| URL | https://arxiv.org/abs/2601.06827 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Detecting pre-training data in Large Language Models (LLMs) is crucial for auditing data privacy and copyright compliance, yet it remains challenging in black-box, zero-shot settings where computational resources and training data are scarce. While existing likelihood-based methods have shown promise, they typically aggregate token-level scores using uniform weights, thereby neglecting the inherent information-theoretic dynamics of autoregressive generation. In this paper, we hypothesize and empirically validate that memorization signals are heavily skewed towards the high-entropy initial tokens, where model uncertainty is highest, and decay as context accumulates. To leverage this linguistic property, we introduce Positional Decay Reweighting (PDR), a training-free and plug-and-play framework. PDR explicitly reweights token-level scores to amplify distinct signals from early positions while suppressing noise from later ones. Extensive experiments show that PDR acts as a robust prior and can usually enhance a wide range of advanced methods across multiple benchmarks.
{
"annotation_id": "af4abf3a-37c6-4e1c-a287-3fef9fd8008b",
"date_created": "2026-02-17T05:53:08.869000Z",
"date_modified": "2026-02-17T05:53:08.869000Z",
"file_hash": "7b8efd5d3b150b944d9e6a0742cd518239469f76e88f294131178bccc6393fc4",
"private": false,
"record": {
"abstract": "Detecting pre-training data in Large Language Models (LLMs) is crucial for auditing data privacy and copyright compliance, yet it remains challenging in black-box, zero-shot settings where computational resources and training data are scarce. While existing likelihood-based methods have shown promise, they typically aggregate token-level scores using uniform weights, thereby neglecting the inherent information-theoretic dynamics of autoregressive generation. In this paper, we hypothesize and empirically validate that memorization signals are heavily skewed towards the high-entropy initial tokens, where model uncertainty is highest, and decay as context accumulates. To leverage this linguistic property, we introduce Positional Decay Reweighting (PDR), a training-free and plug-and-play framework. PDR explicitly reweights token-level scores to amplify distinct signals from early positions while suppressing noise from later ones. Extensive experiments show that PDR acts as a robust prior and can usually enhance a wide range of advanced methods across multiple benchmarks.",
"arxiv_id": "2601.06827",
"authors": [
"Jinhan Liu",
"Yibo Yang",
"Ruiying Lu",
"Piotr Piekos",
"Yimeng Chen",
"Peng Wang",
"Dandan Guo"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "PDR: A Plug-and-Play Positional Decay Framework for LLM Pre-training Data Detection",
"url": "https://arxiv.org/abs/2601.06827",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "eb198a4f-6393-49cc-85d2-a0dda46ef6e6",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}