dorsal/arxiv
View SchemaConvergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-sided and Two-Sided Preconditioning
| Authors | Huan Li, Yiming Dong, Zhouchen Lin |
|---|---|
| Categories | |
| ArXiv ID | 2601.07326vv1 |
| URL | https://arxiv.org/abs/2601.07326 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
This paper studies the AdamW-style Shampoo optimizer, an effective implementation of classical Shampoo that notably won the external tuning track of the AlgoPerf neural network training algorithm competition. Our analysis unifies one-sided and two-sided preconditioning and establishes the convergence rate $\frac{1}{K}\sum_{k=1}^K E\left[\|\nabla f(X_k)\|_*\right]\leq O(\frac{\sqrt{m+n}C}{K^{1/4}})$ measured by nuclear norm, where $K$ represents the iteration number, $(m,n)$ denotes the size of matrix parameters, and $C$ matches the constant in the optimal convergence rate of SGD. Theoretically, we have $\|\nabla f(X)\|_F\leq \|\nabla f(X)\|_*\leq \sqrt{m+n}\|\nabla f(X)\|_F$, supporting that our convergence rate can be considered to be analogous to the optimal $\frac{1}{K}\sum_{k=1}^KE\left[\|\nabla f(X_k)\|_F\right]\leq O(\frac{C}{K^{1/4}})$ convergence rate of SGD in the ideal case of $\|\nabla f(X)\|_*= \Theta(\sqrt{m+n})\|\nabla f(X)\|_F$.
{
"annotation_id": "d0dd66ab-708f-43e8-9d1a-e43a2ff40458",
"date_created": "2026-02-17T05:53:12.436000Z",
"date_modified": "2026-02-17T05:53:12.436000Z",
"file_hash": "83d423eee2715a61c0ed8569764806375da7973ed27e2ceefbefcb58306880c4",
"private": false,
"record": {
"abstract": "This paper studies the AdamW-style Shampoo optimizer, an effective implementation of classical Shampoo that notably won the external tuning track of the AlgoPerf neural network training algorithm competition. Our analysis unifies one-sided and two-sided preconditioning and establishes the convergence rate $\\frac{1}{K}\\sum_{k=1}^K E\\left[\\|\\nabla f(X_k)\\|_*\\right]\\leq O(\\frac{\\sqrt{m+n}C}{K^{1/4}})$ measured by nuclear norm, where $K$ represents the iteration number, $(m,n)$ denotes the size of matrix parameters, and $C$ matches the constant in the optimal convergence rate of SGD. Theoretically, we have $\\|\\nabla f(X)\\|_F\\leq \\|\\nabla f(X)\\|_*\\leq \\sqrt{m+n}\\|\\nabla f(X)\\|_F$, supporting that our convergence rate can be considered to be analogous to the optimal $\\frac{1}{K}\\sum_{k=1}^KE\\left[\\|\\nabla f(X_k)\\|_F\\right]\\leq O(\\frac{C}{K^{1/4}})$ convergence rate of SGD in the ideal case of $\\|\\nabla f(X)\\|_*= \\Theta(\\sqrt{m+n})\\|\\nabla f(X)\\|_F$.",
"arxiv_id": "2601.07326",
"authors": [
"Huan Li",
"Yiming Dong",
"Zhouchen Lin"
],
"categories": [
"math.OC",
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Convergence Rate Analysis of the AdamW-Style Shampoo: Unifying One-sided and Two-Sided Preconditioning",
"url": "https://arxiv.org/abs/2601.07326",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e7fd5b3c-9bf2-4af3-923d-e2d8b609d24b",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}