dorsal/arxiv
View SchemaAPEX: Learning Adaptive Priorities for Multi-Objective Alignment in Vision-Language Generation
| Authors | Dongliang Chen, Xinlin Zhuang, Junjie Xu, Luojian Xie, Zehui Wang, Jiaxi Zhuang, Haolin Yang, Liang Dou, Xiao He, Xingjiao Wu, Ying Qian |
|---|---|
| Categories | |
| ArXiv ID | 2601.06574vv1 |
| URL | https://arxiv.org/abs/2601.06574 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Multi-objective alignment for text-to-image generation is commonly implemented via static linear scalarization, but fixed weights often fail under heterogeneous rewards, leading to optimization imbalance where models overfit high-variance, high-responsiveness objectives (e.g., OCR) while under-optimizing perceptual goals. We identify two mechanistic causes: variance hijacking, where reward dispersion induces implicit reweighting that dominates the normalized training signal, and gradient conflicts, where competing objectives produce opposing update directions and trigger seesaw-like oscillations. We propose APEX (Adaptive Priority-based Efficient X-objective Alignment), which stabilizes heterogeneous rewards with Dual-Stage Adaptive Normalization and dynamically schedules objectives via P^3 Adaptive Priorities that combine learning potential, conflict penalty, and progress need. On Stable Diffusion 3.5, APEX achieves improved Pareto trade-offs across four heterogeneous objectives, with balanced gains of +1.31 PickScore, +0.35 DeQA, and +0.53 Aesthetics while maintaining competitive OCR accuracy, mitigating the instability of multi-objective alignment.
{
"annotation_id": "98389197-eef3-4585-a616-3e4486d84659",
"date_created": "2026-02-17T05:53:07.771000Z",
"date_modified": "2026-02-17T05:53:07.771000Z",
"file_hash": "bc3423a4c0ee1b9ea5f479cd8d39ffa83e37921d52825def8f8506a2d09ad185",
"private": false,
"record": {
"abstract": "Multi-objective alignment for text-to-image generation is commonly implemented via static linear scalarization, but fixed weights often fail under heterogeneous rewards, leading to optimization imbalance where models overfit high-variance, high-responsiveness objectives (e.g., OCR) while under-optimizing perceptual goals. We identify two mechanistic causes: variance hijacking, where reward dispersion induces implicit reweighting that dominates the normalized training signal, and gradient conflicts, where competing objectives produce opposing update directions and trigger seesaw-like oscillations. We propose APEX (Adaptive Priority-based Efficient X-objective Alignment), which stabilizes heterogeneous rewards with Dual-Stage Adaptive Normalization and dynamically schedules objectives via P^3 Adaptive Priorities that combine learning potential, conflict penalty, and progress need. On Stable Diffusion 3.5, APEX achieves improved Pareto trade-offs across four heterogeneous objectives, with balanced gains of +1.31 PickScore, +0.35 DeQA, and +0.53 Aesthetics while maintaining competitive OCR accuracy, mitigating the instability of multi-objective alignment.",
"arxiv_id": "2601.06574",
"authors": [
"Dongliang Chen",
"Xinlin Zhuang",
"Junjie Xu",
"Luojian Xie",
"Zehui Wang",
"Jiaxi Zhuang",
"Haolin Yang",
"Liang Dou",
"Xiao He",
"Xingjiao Wu",
"Ying Qian"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "APEX: Learning Adaptive Priorities for Multi-Objective Alignment in Vision-Language Generation",
"url": "https://arxiv.org/abs/2601.06574",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "bc903fce-baac-4844-ab04-f947cd75709a",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}