dorsal/arxiv
View SchemaControlled LLM Training on Spectral Sphere
| Authors | Tian Xie, Haoming Luo, Haoyu Tang, Yiwen Hu, Jason Klein Liu, Qingnan Ren, Yang Wang, Wayne Xin Zhao, Rui Yan, Bing Su, Chong Luo, Baining Guo |
|---|---|
| Categories | |
| ArXiv ID | 2601.08393vv2 |
| URL | https://arxiv.org/abs/2601.08393 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization ($\boldsymbol{\mu}$P) provides a theoretical safeguard for width-invariant $\Theta(1)$ activation control, whereas emerging optimizers like Muon are only ``half-aligned'' with these constraints: they control updates but allow weights to drift. To address this limitation, we introduce the \textbf{Spectral Sphere Optimizer (SSO)}, which enforces strict module-wise spectral constraints on both weights and their updates. By deriving the steepest descent direction on the spectral sphere, SSO realizes a fully $\boldsymbol{\mu}$P-aligned optimization process. To enable large-scale training, we implement SSO as an efficient parallel algorithm within Megatron. Through extensive pretraining on diverse architectures, including Dense 1.7B, MoE 8B-A1B, and 200-layer DeepNet models, SSO consistently outperforms AdamW and Muon. Furthermore, we observe significant practical stability benefits, including improved MoE router load balancing, suppressed outliers, and strictly bounded activations.
{
"annotation_id": "42657afe-c63d-4f60-861d-a2c0ee61d820",
"date_created": "2026-02-17T05:53:15.835000Z",
"date_modified": "2026-02-17T05:53:15.835000Z",
"file_hash": "9d3213cd004120447db7f8b3c4b3c4c368a2fd79cd9fb3995ef65e523c1ddc0f",
"private": false,
"record": {
"abstract": "Scaling large models requires optimization strategies that ensure rapid convergence grounded in stability. Maximal Update Parametrization ($\\boldsymbol{\\mu}$P) provides a theoretical safeguard for width-invariant $\\Theta(1)$ activation control, whereas emerging optimizers like Muon are only ``half-aligned\u0027\u0027 with these constraints: they control updates but allow weights to drift. To address this limitation, we introduce the \\textbf{Spectral Sphere Optimizer (SSO)}, which enforces strict module-wise spectral constraints on both weights and their updates. By deriving the steepest descent direction on the spectral sphere, SSO realizes a fully $\\boldsymbol{\\mu}$P-aligned optimization process. To enable large-scale training, we implement SSO as an efficient parallel algorithm within Megatron. Through extensive pretraining on diverse architectures, including Dense 1.7B, MoE 8B-A1B, and 200-layer DeepNet models, SSO consistently outperforms AdamW and Muon. Furthermore, we observe significant practical stability benefits, including improved MoE router load balancing, suppressed outliers, and strictly bounded activations.",
"arxiv_id": "2601.08393",
"authors": [
"Tian Xie",
"Haoming Luo",
"Haoyu Tang",
"Yiwen Hu",
"Jason Klein Liu",
"Qingnan Ren",
"Yang Wang",
"Wayne Xin Zhao",
"Rui Yan",
"Bing Su",
"Chong Luo",
"Baining Guo"
],
"categories": [
"cs.LG",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Controlled LLM Training on Spectral Sphere",
"url": "https://arxiv.org/abs/2601.08393",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "61114556-fbe4-4951-972b-8e1a5c80df01",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}