dorsal/arxiv
View SchemaSTAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio
| Authors | Seong-Gyu Park, Sohee Park, Jisu Lee, Hyunsik Na, Daeseon Choi |
|---|---|
| Categories | |
| ArXiv ID | 2601.08511vv1 |
| URL | https://arxiv.org/abs/2601.08511 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Recent LLMs increasingly integrate reasoning mechanisms like Chain-of-Thought (CoT). However, this explicit reasoning exposes a new attack surface for inference-time backdoors, which inject malicious reasoning paths without altering model parameters. Because these attacks generate linguistically coherent paths, they effectively evade conventional detection. To address this, we propose STAR (State-Transition Amplification Ratio), a framework that detects backdoors by analyzing output probability shifts. STAR exploits the statistical discrepancy where a malicious input-induced path exhibits high posterior probability despite a low prior probability in the model's general knowledge. We quantify this state-transition amplification and employ the CUSUM algorithm to detect persistent anomalies. Experiments across diverse models (8B-70B) and five benchmark datasets demonstrate that STAR exhibits robust generalization capabilities, consistently achieving near-perfect performance (AUROC $\approx$ 1.0) with approximately $42\times$ greater efficiency than existing baselines. Furthermore, the framework proves robust against adaptive attacks attempting to bypass detection.
{
"annotation_id": "4c7577e0-5586-4ccf-ae50-5251adfed83a",
"date_created": "2026-02-17T05:53:15.623000Z",
"date_modified": "2026-02-17T05:53:15.623000Z",
"file_hash": "d5408287e82bfa730bbde3473d661b8f748dee4ef7fb4087c4ed00789ca9cbf4",
"private": false,
"record": {
"abstract": "Recent LLMs increasingly integrate reasoning mechanisms like Chain-of-Thought (CoT). However, this explicit reasoning exposes a new attack surface for inference-time backdoors, which inject malicious reasoning paths without altering model parameters. Because these attacks generate linguistically coherent paths, they effectively evade conventional detection. To address this, we propose STAR (State-Transition Amplification Ratio), a framework that detects backdoors by analyzing output probability shifts. STAR exploits the statistical discrepancy where a malicious input-induced path exhibits high posterior probability despite a low prior probability in the model\u0027s general knowledge. We quantify this state-transition amplification and employ the CUSUM algorithm to detect persistent anomalies. Experiments across diverse models (8B-70B) and five benchmark datasets demonstrate that STAR exhibits robust generalization capabilities, consistently achieving near-perfect performance (AUROC $\\approx$ 1.0) with approximately $42\\times$ greater efficiency than existing baselines. Furthermore, the framework proves robust against adaptive attacks attempting to bypass detection.",
"arxiv_id": "2601.08511",
"authors": [
"Seong-Gyu Park",
"Sohee Park",
"Jisu Lee",
"Hyunsik Na",
"Daeseon Choi"
],
"categories": [
"cs.CL",
"cs.CR",
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio",
"url": "https://arxiv.org/abs/2601.08511",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "4ba9f4dc-df0b-434e-90bd-70aafda3922d",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}