dorsal/arxiv
View SchemaLDLT L-Lipschitz Network Weight Parameterization Initialization
| Authors | Marius F. R. Juston, Ramavarapu S. Sreenivas, Dustin Nottage, Ahmet Soylemezoglu |
|---|---|
| Categories | |
| ArXiv ID | 2601.08253vv1 |
| URL | https://arxiv.org/abs/2601.08253 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
We analyze initialization dynamics for LDLT-based $\mathcal{L}$-Lipschitz layers by deriving the exact marginal output variance when the underlying parameter matrix $W_0\in \mathbb{R}^{m\times n}$ is initialized with IID Gaussian entries $\mathcal{N}(0,\sigma^2)$. The Wishart distribution, $S=W_0W_0^\top\sim\mathcal{W}_m(n,\sigma^2 \boldsymbol{I}_m)$, used for computing the output marginal variance is derived in closed form using expectations of zonal polynomials via James' theorem and a Laplace-integral expansion of $(\alpha \boldsymbol{I}_m+S)^{-1}$. We develop an Isserlis/Wick-based combinatorial expansion for $\operatorname{\mathbb{E}}\left[\operatorname{tr}(S^k)\right]$ and provide explicit truncated moments up to $k=10$, which yield accurate series approximations for small-to-moderate $\sigma^2$. Monte Carlo experiments confirm the theoretical estimates. Furthermore, empirical analysis was performed to quantify that, using current He or Kaiming initialization with scaling $1/\sqrt{n}$, the output variance is $0.41$, whereas the new parameterization with $10/ \sqrt{n}$ for $\alpha=1$ results in an output variance of $0.9$. The findings clarify why deep $\mathcal{L}$-Lipschitz networks suffer rapid information loss at initialization and offer practical prescriptions for choosing initialization hyperparameters to mitigate this effect. However, using the Higgs boson classification dataset, a hyperparameter sweep over optimizers, initialization scale, and depth was conducted to validate the results on real-world data, showing that although the derivation ensures variance preservation, empirical results indicate He initialization still performs better.
{
"annotation_id": "290d4726-48a0-4faa-9bbd-0eda18c3e645",
"date_created": "2026-02-17T05:53:15.437000Z",
"date_modified": "2026-02-17T05:53:15.437000Z",
"file_hash": "3cdeb5da5fab478e04f129ab213399aae548201a712a3a5b26ccb29037a9e51d",
"private": false,
"record": {
"abstract": "We analyze initialization dynamics for LDLT-based $\\mathcal{L}$-Lipschitz layers by deriving the exact marginal output variance when the underlying parameter matrix $W_0\\in \\mathbb{R}^{m\\times n}$ is initialized with IID Gaussian entries $\\mathcal{N}(0,\\sigma^2)$. The Wishart distribution, $S=W_0W_0^\\top\\sim\\mathcal{W}_m(n,\\sigma^2 \\boldsymbol{I}_m)$, used for computing the output marginal variance is derived in closed form using expectations of zonal polynomials via James\u0027 theorem and a Laplace-integral expansion of $(\\alpha \\boldsymbol{I}_m+S)^{-1}$. We develop an Isserlis/Wick-based combinatorial expansion for $\\operatorname{\\mathbb{E}}\\left[\\operatorname{tr}(S^k)\\right]$ and provide explicit truncated moments up to $k=10$, which yield accurate series approximations for small-to-moderate $\\sigma^2$. Monte Carlo experiments confirm the theoretical estimates. Furthermore, empirical analysis was performed to quantify that, using current He or Kaiming initialization with scaling $1/\\sqrt{n}$, the output variance is $0.41$, whereas the new parameterization with $10/ \\sqrt{n}$ for $\\alpha=1$ results in an output variance of $0.9$. The findings clarify why deep $\\mathcal{L}$-Lipschitz networks suffer rapid information loss at initialization and offer practical prescriptions for choosing initialization hyperparameters to mitigate this effect. However, using the Higgs boson classification dataset, a hyperparameter sweep over optimizers, initialization scale, and depth was conducted to validate the results on real-world data, showing that although the derivation ensures variance preservation, empirical results indicate He initialization still performs better.",
"arxiv_id": "2601.08253",
"authors": [
"Marius F. R. Juston",
"Ramavarapu S. Sreenivas",
"Dustin Nottage",
"Ahmet Soylemezoglu"
],
"categories": [
"cs.LG"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "LDLT L-Lipschitz Network Weight Parameterization Initialization",
"url": "https://arxiv.org/abs/2601.08253",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "0637f06e-68ae-4d93-a4e1-ead2bd2a1df0",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}