dorsal/arxiv
View SchemaDistribution Estimation with Side Information
| Authors | Haricharan Balasundaram, Andrew Thangaraj |
|---|---|
| Categories | |
| ArXiv ID | 2601.08535vv2 |
| URL | https://arxiv.org/abs/2601.08535 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side information arises naturally through word semantics/similarities that can be inferred by closeness of vector word embeddings, for instance. We consider two specific models for side information--a local model where the unknown distribution is in the neighborhood of a known distribution, and a partial ordering model where the alphabet is partitioned into known higher and lower probability sets. In both models, we theoretically characterize the improvement in a suitable squared-error risk because of the available side information. Simulations over natural language and synthetic data illustrate these gains.
{
"annotation_id": "4c343a12-4f5b-4800-9d4d-b16459c8b6b2",
"date_created": "2026-02-17T05:53:16.126000Z",
"date_modified": "2026-02-17T05:53:16.126000Z",
"file_hash": "01279189fb18c3e9c6f939399761ebbbc20724efbe79a7a2ef00dc70e493182b",
"private": false,
"record": {
"abstract": "We consider the classical problem of discrete distribution estimation using i.i.d. samples in a novel scenario where additional side information is available on the distribution. In large alphabet datasets such as text corpora, such side information arises naturally through word semantics/similarities that can be inferred by closeness of vector word embeddings, for instance. We consider two specific models for side information--a local model where the unknown distribution is in the neighborhood of a known distribution, and a partial ordering model where the alphabet is partitioned into known higher and lower probability sets. In both models, we theoretically characterize the improvement in a suitable squared-error risk because of the available side information. Simulations over natural language and synthetic data illustrate these gains.",
"arxiv_id": "2601.08535",
"authors": [
"Haricharan Balasundaram",
"Andrew Thangaraj"
],
"categories": [
"cs.IT",
"math.IT"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Distribution Estimation with Side Information",
"url": "https://arxiv.org/abs/2601.08535",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "4cfd28cc-c3dc-4312-b1d1-173235ee8053",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}