dorsal/arxiv
View SchemaHellinger Multimodal Variational Autoencoders
| Authors | Huyen Khanh Vo, Isabel Valera |
|---|---|
| Categories | |
| ArXiv ID | 2601.06572vv1 |
| URL | https://arxiv.org/abs/2601.06572 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities. Predominant methods aggregate unimodal inference distributions using either a product of experts (PoE), a mixture of experts (MoE), or their combinations to approximate the joint posterior. In this work, we revisit multimodal inference through the lens of probabilistic opinion pooling, an optimization-based approach. We start from H\"older pooling with $\alpha=0.5$, which corresponds to the unique symmetric member of the $\alpha\text{-divergence}$ family, and derive a moment-matching approximation, termed Hellinger. We then leverage such an approximation to propose HELVAE, a multimodal VAE that avoids sub-sampling, yielding an efficient yet effective model that: (i) learns more expressive latent representations as additional modalities are observed; and (ii) empirically achieves better trade-offs between generative coherence and quality, outperforming state-of-the-art multimodal VAE models.
{
"annotation_id": "12feea47-3a02-44cf-bb3b-5862ee077757",
"date_created": "2026-02-17T05:53:07.773000Z",
"date_modified": "2026-02-17T05:53:07.773000Z",
"file_hash": "0a6b2a91d8f3ed61da6d03908dea6ef51ee193e46fdcfd47e57898fa3c4494de",
"private": false,
"record": {
"abstract": "Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities. Predominant methods aggregate unimodal inference distributions using either a product of experts (PoE), a mixture of experts (MoE), or their combinations to approximate the joint posterior. In this work, we revisit multimodal inference through the lens of probabilistic opinion pooling, an optimization-based approach. We start from H\\\"older pooling with $\\alpha=0.5$, which corresponds to the unique symmetric member of the $\\alpha\\text{-divergence}$ family, and derive a moment-matching approximation, termed Hellinger. We then leverage such an approximation to propose HELVAE, a multimodal VAE that avoids sub-sampling, yielding an efficient yet effective model that: (i) learns more expressive latent representations as additional modalities are observed; and (ii) empirically achieves better trade-offs between generative coherence and quality, outperforming state-of-the-art multimodal VAE models.",
"arxiv_id": "2601.06572",
"authors": [
"Huyen Khanh Vo",
"Isabel Valera"
],
"categories": [
"cs.LG",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Hellinger Multimodal Variational Autoencoders",
"url": "https://arxiv.org/abs/2601.06572",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "70cead1e-1e18-46c3-805c-983b50df9ca4",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}