dorsal/arxiv
View SchemaSteer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding
| Authors | Yijiang River Dong, Tiancheng Hu, Zheng Hui, Nigel Collier |
|---|---|
| Categories | |
| ArXiv ID | 2601.06403vv1 |
| URL | https://arxiv.org/abs/2601.06403 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large language models excel at complex instructions yet struggle to deviate from their helpful assistant persona, as post-training instills strong priors that resist conflicting instructions. We introduce system prompt strength, a training-free method that treats prompt adherence as a continuous control. By contrasting logits from target and default system prompts, we isolate and amplify the behavioral signal unique to the target persona by a scalar factor alpha. Across five diverse benchmarks spanning constraint satisfaction, behavioral control, pluralistic alignment, capability modulation, and stylistic control, our method yields substantial improvements: up to +8.5 strict accuracy on IFEval, +45pp refusal rate on OffTopicEval, and +13% steerability on Prompt-Steering. Our approach enables practitioners to modulate system prompt strength, providing dynamic control over model behavior without retraining.
{
"annotation_id": "ff7092e1-601a-4c7b-a84d-e4393edfcfec",
"date_created": "2026-02-17T05:53:08.236000Z",
"date_modified": "2026-02-17T05:53:08.236000Z",
"file_hash": "f4582ca37bf68ea5de24c81004c96a9fd3e3e57225f41654368f7ecbed3034f4",
"private": false,
"record": {
"abstract": "Large language models excel at complex instructions yet struggle to deviate from their helpful assistant persona, as post-training instills strong priors that resist conflicting instructions. We introduce system prompt strength, a training-free method that treats prompt adherence as a continuous control. By contrasting logits from target and default system prompts, we isolate and amplify the behavioral signal unique to the target persona by a scalar factor alpha. Across five diverse benchmarks spanning constraint satisfaction, behavioral control, pluralistic alignment, capability modulation, and stylistic control, our method yields substantial improvements: up to +8.5 strict accuracy on IFEval, +45pp refusal rate on OffTopicEval, and +13% steerability on Prompt-Steering. Our approach enables practitioners to modulate system prompt strength, providing dynamic control over model behavior without retraining.",
"arxiv_id": "2601.06403",
"authors": [
"Yijiang River Dong",
"Tiancheng Hu",
"Zheng Hui",
"Nigel Collier"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Steer Model beyond Assistant: Controlling System Prompt Strength via Contrastive Decoding",
"url": "https://arxiv.org/abs/2601.06403",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "25f043ab-5b7a-47dc-8dd7-afa0c038a2fb",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}