dorsal/arxiv
View SchemaCoverage Improvement and Fast Convergence of On-policy Preference Learning
| Authors | Juno Kim, Jihun Yun, Jason D. Lee, Kwang-Sung Jun |
|---|---|
| Categories | |
| ArXiv ID | 2601.08421vv1 |
| URL | https://arxiv.org/abs/2601.08421 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Online on-policy preference learning algorithms for language model alignment such as online direct policy optimization (DPO) can significantly outperform their offline counterparts. We provide a theoretical explanation for this phenomenon by analyzing how the sampling policy's coverage evolves throughout on-policy training. We propose and rigorously justify the \emph{coverage improvement principle}: with sufficient batch size, each update moves into a region around the target where coverage is uniformly better, making subsequent data increasingly informative and enabling rapid convergence. In the contextual bandit setting with Bradley-Terry preferences and linear softmax policy class, we show that on-policy DPO converges exponentially in the number of iterations for batch size exceeding a generalized coverage threshold. In contrast, any learner restricted to offline samples from the initial policy suffers a slower minimax rate, leading to a sharp separation in total sample complexity. Motivated by this analysis, we further propose a simple hybrid sampler based on a novel \emph{preferential} G-optimal design, which removes dependence on coverage and guarantees convergence in just two rounds. Finally, we develop principled on-policy schemes for reward distillation in the general function class setting, and show faster noiseless rates under an alternative deviation-based notion of coverage. Experimentally, we confirm that on-policy DPO and our proposed reward distillation algorithms outperform their off-policy counterparts and enjoy stable, monotonic performance gains across iterations.
{
"annotation_id": "bf95d691-ca0c-4a87-a17e-5d7e3f8faf94",
"date_created": "2026-02-17T05:53:16.200000Z",
"date_modified": "2026-02-17T05:53:16.200000Z",
"file_hash": "0e1812ac08dd167ef2be70f0af06e9c7f1cb4e064f71a77c06e1b43d13468abd",
"private": false,
"record": {
"abstract": "Online on-policy preference learning algorithms for language model alignment such as online direct policy optimization (DPO) can significantly outperform their offline counterparts. We provide a theoretical explanation for this phenomenon by analyzing how the sampling policy\u0027s coverage evolves throughout on-policy training. We propose and rigorously justify the \\emph{coverage improvement principle}: with sufficient batch size, each update moves into a region around the target where coverage is uniformly better, making subsequent data increasingly informative and enabling rapid convergence. In the contextual bandit setting with Bradley-Terry preferences and linear softmax policy class, we show that on-policy DPO converges exponentially in the number of iterations for batch size exceeding a generalized coverage threshold. In contrast, any learner restricted to offline samples from the initial policy suffers a slower minimax rate, leading to a sharp separation in total sample complexity. Motivated by this analysis, we further propose a simple hybrid sampler based on a novel \\emph{preferential} G-optimal design, which removes dependence on coverage and guarantees convergence in just two rounds. Finally, we develop principled on-policy schemes for reward distillation in the general function class setting, and show faster noiseless rates under an alternative deviation-based notion of coverage. Experimentally, we confirm that on-policy DPO and our proposed reward distillation algorithms outperform their off-policy counterparts and enjoy stable, monotonic performance gains across iterations.",
"arxiv_id": "2601.08421",
"authors": [
"Juno Kim",
"Jihun Yun",
"Jason D. Lee",
"Kwang-Sung Jun"
],
"categories": [
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Coverage Improvement and Fast Convergence of On-policy Preference Learning",
"url": "https://arxiv.org/abs/2601.08421",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6c062a5d-f359-4939-94d1-4da2213a79a8",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}