dorsal/arxiv
View SchemaGIFT: Unlocking Global Optimality in Post-Training via Finite-Temperature Gibbs Initialization
| Authors | Zhengyang Zhao, Lu Ma, Yizhen Jiang, Xiaochen Ma, Zimo Meng, Chengyu Shen, Lexiang Tang, Haoze Sun, Peng Pei, Wentao Zhang |
|---|---|
| Categories | |
| ArXiv ID | 2601.09233vv1 |
| URL | https://arxiv.org/abs/2601.09233 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
The prevailing post-training paradigm for Large Reasoning Models (LRMs)--Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL)--suffers from an intrinsic optimization mismatch: the rigid supervision inherent in SFT induces distributional collapse, thereby exhausting the exploration space necessary for subsequent RL. In this paper, we reformulate SFT within a unified post-training framework and propose Gibbs Initialization with Finite Temperature (GIFT). We characterize standard SFT as a degenerate zero-temperature limit that suppresses base priors. Conversely, GIFT incorporates supervision as a finite-temperature energy potential, establishing a distributional bridge that ensures objective consistency throughout the post-training pipeline. Our experiments demonstrate that GIFT significantly outperforms standard SFT and other competitive baselines when utilized for RL initialization, providing a mathematically principled pathway toward achieving global optimality in post-training. Our code is available at https://github.com/zzy1127/GIFT.
{
"annotation_id": "3956ea79-3e36-4215-9b70-62dddae7a025",
"date_created": "2026-02-17T05:53:20.215000Z",
"date_modified": "2026-02-17T05:53:20.215000Z",
"file_hash": "8d8252caba427605b695c814ad48f07ff754433dd32f42dc0b879f748f5a2c57",
"private": false,
"record": {
"abstract": "The prevailing post-training paradigm for Large Reasoning Models (LRMs)--Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL)--suffers from an intrinsic optimization mismatch: the rigid supervision inherent in SFT induces distributional collapse, thereby exhausting the exploration space necessary for subsequent RL. In this paper, we reformulate SFT within a unified post-training framework and propose Gibbs Initialization with Finite Temperature (GIFT). We characterize standard SFT as a degenerate zero-temperature limit that suppresses base priors. Conversely, GIFT incorporates supervision as a finite-temperature energy potential, establishing a distributional bridge that ensures objective consistency throughout the post-training pipeline. Our experiments demonstrate that GIFT significantly outperforms standard SFT and other competitive baselines when utilized for RL initialization, providing a mathematically principled pathway toward achieving global optimality in post-training. Our code is available at https://github.com/zzy1127/GIFT.",
"arxiv_id": "2601.09233",
"authors": [
"Zhengyang Zhao",
"Lu Ma",
"Yizhen Jiang",
"Xiaochen Ma",
"Zimo Meng",
"Chengyu Shen",
"Lexiang Tang",
"Haoze Sun",
"Peng Pei",
"Wentao Zhang"
],
"categories": [
"cs.LG",
"cs.AI",
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "GIFT: Unlocking Global Optimality in Post-Training via Finite-Temperature Gibbs Initialization",
"url": "https://arxiv.org/abs/2601.09233",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "78e02370-4aea-4074-8994-99f798df0033",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}