dorsal/arxiv
View SchemaMid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers
| Authors | Wang Yang, Debargha Ganguly, Xinpeng Li, Chaoda Song, Shouren Wang, Vikash Singh, Vipin Chaudhary, Xiaotian Han |
|---|---|
| Categories | |
| ArXiv ID | 2601.07036vv1 |
| URL | https://arxiv.org/abs/2601.07036 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Hybrid reasoning language models are commonly controlled through high-level Think/No-think instructions to regulate reasoning behavior, yet we found that such mode switching is largely driven by a small set of trigger tokens rather than the instructions themselves. Through attention analysis and controlled prompting experiments, we show that a leading ``Okay'' token induces reasoning behavior, while the newline pattern following ``</think>'' suppresses it. Based on this observation, we propose Mid-Think, a simple training-free prompting format that combines these triggers to achieve intermediate-budget reasoning, consistently outperforming fixed-token and prompt-based baselines in terms of the accuracy-length trade-off. Furthermore, applying Mid-Think to RL training after SFT reduces training time by approximately 15% while improving final performance of Qwen3-8B on AIME from 69.8% to 72.4% and on GPQA from 58.5% to 61.1%, demonstrating its effectiveness for both inference-time control and RL-based reasoning training.
{
"annotation_id": "22c4598c-314a-40e5-8bff-97073b8440a0",
"date_created": "2026-02-17T05:53:08.870000Z",
"date_modified": "2026-02-17T05:53:08.870000Z",
"file_hash": "9d245c94bfa64e96fee5e554edbbc8b7528ed5526e324b0a494a0b4f151f2a8f",
"private": false,
"record": {
"abstract": "Hybrid reasoning language models are commonly controlled through high-level Think/No-think instructions to regulate reasoning behavior, yet we found that such mode switching is largely driven by a small set of trigger tokens rather than the instructions themselves. Through attention analysis and controlled prompting experiments, we show that a leading ``Okay\u0027\u0027 token induces reasoning behavior, while the newline pattern following ``\u003c/think\u003e\u0027\u0027 suppresses it. Based on this observation, we propose Mid-Think, a simple training-free prompting format that combines these triggers to achieve intermediate-budget reasoning, consistently outperforming fixed-token and prompt-based baselines in terms of the accuracy-length trade-off. Furthermore, applying Mid-Think to RL training after SFT reduces training time by approximately 15% while improving final performance of Qwen3-8B on AIME from 69.8% to 72.4% and on GPQA from 58.5% to 61.1%, demonstrating its effectiveness for both inference-time control and RL-based reasoning training.",
"arxiv_id": "2601.07036",
"authors": [
"Wang Yang",
"Debargha Ganguly",
"Xinpeng Li",
"Chaoda Song",
"Shouren Wang",
"Vikash Singh",
"Vipin Chaudhary",
"Xiaotian Han"
],
"categories": [
"cs.CL",
"cs.AI",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Mid-Think: Training-Free Intermediate-Budget Reasoning via Token-Level Triggers",
"url": "https://arxiv.org/abs/2601.07036",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6e614264-7198-45a2-adc8-d9bb198fb006",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}