dorsal/arxiv
View SchemaDeriving Decoder-Free Sparse Autoencoders from First Principles
| Authors | Alan Oursland |
|---|---|
| Categories | |
| ArXiv ID | 2601.06478vv1 |
| URL | https://arxiv.org/abs/2601.06478 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Gradient descent on log-sum-exp (LSE) objectives performs implicit expectation--maximization (EM): the gradient with respect to each component output equals its responsibility. The same theory predicts collapse without volume control analogous to the log-determinant in Gaussian mixture models. We instantiate the theory in a single-layer encoder with an LSE objective and InfoMax regularization for volume control. Experiments confirm the theory's predictions. The gradient--responsibility identity holds exactly; LSE alone collapses; variance prevents dead components; decorrelation prevents redundancy. The model exhibits EM-like optimization dynamics in which lower loss does not correspond to better features and adaptive optimizers offer no advantage. The resulting decoder-free model learns interpretable mixture components, confirming that implicit EM theory can prescribe architectures.
{
"annotation_id": "a7634191-becf-4567-8ecc-7a5e703ef113",
"date_created": "2026-02-17T05:53:08.036000Z",
"date_modified": "2026-02-17T05:53:08.036000Z",
"file_hash": "37b994f8e6f6bb3d879bb881e3e15fb24d36e6d26c8840cc6df3e1e48136508f",
"private": false,
"record": {
"abstract": "Gradient descent on log-sum-exp (LSE) objectives performs implicit expectation--maximization (EM): the gradient with respect to each component output equals its responsibility. The same theory predicts collapse without volume control analogous to the log-determinant in Gaussian mixture models. We instantiate the theory in a single-layer encoder with an LSE objective and InfoMax regularization for volume control. Experiments confirm the theory\u0027s predictions. The gradient--responsibility identity holds exactly; LSE alone collapses; variance prevents dead components; decorrelation prevents redundancy. The model exhibits EM-like optimization dynamics in which lower loss does not correspond to better features and adaptive optimizers offer no advantage. The resulting decoder-free model learns interpretable mixture components, confirming that implicit EM theory can prescribe architectures.",
"arxiv_id": "2601.06478",
"authors": [
"Alan Oursland"
],
"categories": [
"cs.LG",
"stat.ML"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Deriving Decoder-Free Sparse Autoencoders from First Principles",
"url": "https://arxiv.org/abs/2601.06478",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "ddf8bdde-6617-4182-b7d7-c5002a320dce",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}