dorsal/arxiv
View SchemaReverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies
| Authors | Zeyang Li, Sunbochen Tang, Navid Azizan |
|---|---|
| Categories | |
| ArXiv ID | 2601.08136vv1 |
| URL | https://arxiv.org/abs/2601.08136 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Diffusion and flow policies are gaining prominence in online reinforcement learning (RL) due to their expressive power, yet training them efficiently remains a critical challenge. A fundamental difficulty in online RL is the lack of direct samples from the target distribution; instead, the target is an unnormalized Boltzmann distribution defined by the Q-function. To address this, two seemingly distinct families of methods have been proposed for diffusion policies: a noise-expectation family, which utilizes a weighted average of noise as the training target, and a gradient-expectation family, which employs a weighted average of Q-function gradients. Yet, it remains unclear how these objectives relate formally or if they can be synthesized into a more general formulation. In this paper, we propose a unified framework, reverse flow matching (RFM), which rigorously addresses the problem of training diffusion and flow models without direct target samples. By adopting a reverse inferential perspective, we formulate the training target as a posterior mean estimation problem given an intermediate noisy sample. Crucially, we introduce Langevin Stein operators to construct zero-mean control variates, deriving a general class of estimators that effectively reduce importance sampling variance. We show that existing noise-expectation and gradient-expectation methods are two specific instances within this broader class. This unified view yields two key advancements: it extends the capability of targeting Boltzmann distributions from diffusion to flow policies, and enables the principled combination of Q-value and Q-gradient information to derive an optimal, minimum-variance estimator, thereby improving training efficiency and stability. We instantiate RFM to train a flow policy in online RL, and demonstrate improved performance on continuous-control benchmarks compared to diffusion policy baselines.
{
"annotation_id": "96428c2e-f949-4436-be07-8ce9cdafb0a3",
"date_created": "2026-02-17T05:53:16.084000Z",
"date_modified": "2026-02-17T05:53:16.084000Z",
"file_hash": "b95c098d5152347c9ff9c604654cf48f727bf81943f0573bd680673eec37ac9a",
"private": false,
"record": {
"abstract": "Diffusion and flow policies are gaining prominence in online reinforcement learning (RL) due to their expressive power, yet training them efficiently remains a critical challenge. A fundamental difficulty in online RL is the lack of direct samples from the target distribution; instead, the target is an unnormalized Boltzmann distribution defined by the Q-function. To address this, two seemingly distinct families of methods have been proposed for diffusion policies: a noise-expectation family, which utilizes a weighted average of noise as the training target, and a gradient-expectation family, which employs a weighted average of Q-function gradients. Yet, it remains unclear how these objectives relate formally or if they can be synthesized into a more general formulation. In this paper, we propose a unified framework, reverse flow matching (RFM), which rigorously addresses the problem of training diffusion and flow models without direct target samples. By adopting a reverse inferential perspective, we formulate the training target as a posterior mean estimation problem given an intermediate noisy sample. Crucially, we introduce Langevin Stein operators to construct zero-mean control variates, deriving a general class of estimators that effectively reduce importance sampling variance. We show that existing noise-expectation and gradient-expectation methods are two specific instances within this broader class. This unified view yields two key advancements: it extends the capability of targeting Boltzmann distributions from diffusion to flow policies, and enables the principled combination of Q-value and Q-gradient information to derive an optimal, minimum-variance estimator, thereby improving training efficiency and stability. We instantiate RFM to train a flow policy in online RL, and demonstrate improved performance on continuous-control benchmarks compared to diffusion policy baselines.",
"arxiv_id": "2601.08136",
"authors": [
"Zeyang Li",
"Sunbochen Tang",
"Navid Azizan"
],
"categories": [
"cs.LG",
"cs.SY",
"eess.SY"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Reverse Flow Matching: A Unified Framework for Online Reinforcement Learning with Diffusion and Flow Policies",
"url": "https://arxiv.org/abs/2601.08136",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6431b8f2-3e28-40a7-b659-77b268098a6c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}