dorsal/arxiv
View SchemaMulti-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation
| Authors | Aaron R. Flouro, Shawn P. Chadwick |
|---|---|
| Categories | |
| ArXiv ID | 2601.09165vv1 |
| URL | https://arxiv.org/abs/2601.09165 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Building on the probability-domain distillation framework of Sparse-KD, we develop an axiomatic, operator-theoretic framework for multi-teacher ensemble knowledge distillation. Rather than prescribing a specific aggregation formula, we define five core axioms governing valid knowledge aggregation operators, encompassing convexity, positivity, continuity, weight monotonicity, and temperature coherence. We prove the existence and non-uniqueness of operator families satisfying these axioms, establishing that multiple distinct aggregation mechanisms conform to the same foundational principles. Within this framework, we establish operator-agnostic guarantees showing that multi-teacher aggregation reduces both stochastic variance and systematic supervisory bias under heterogeneous teachers, while providing Jensen-type bounds, log-loss guarantees, and safety attenuation properties. For aggregation operators linear in teacher weights, we further establish classical ensemble variance-reduction results under standard independence assumptions, with extensions to correlated-error regimes. The framework provides theoretical grounding for multi-teacher distillation from diverse frontier models while admitting multiple valid implementation strategies.
{
"annotation_id": "d08161d4-f7aa-45ab-acb4-dada32a92be5",
"date_created": "2026-02-17T05:53:20.075000Z",
"date_modified": "2026-02-17T05:53:20.075000Z",
"file_hash": "92ee48e02d86b46fcd4d915965ea6351b70f4d865e35ac2569b1d50fa1f3b2f5",
"private": false,
"record": {
"abstract": "Building on the probability-domain distillation framework of Sparse-KD, we develop an axiomatic, operator-theoretic framework for multi-teacher ensemble knowledge distillation. Rather than prescribing a specific aggregation formula, we define five core axioms governing valid knowledge aggregation operators, encompassing convexity, positivity, continuity, weight monotonicity, and temperature coherence. We prove the existence and non-uniqueness of operator families satisfying these axioms, establishing that multiple distinct aggregation mechanisms conform to the same foundational principles.\n Within this framework, we establish operator-agnostic guarantees showing that multi-teacher aggregation reduces both stochastic variance and systematic supervisory bias under heterogeneous teachers, while providing Jensen-type bounds, log-loss guarantees, and safety attenuation properties. For aggregation operators linear in teacher weights, we further establish classical ensemble variance-reduction results under standard independence assumptions, with extensions to correlated-error regimes. The framework provides theoretical grounding for multi-teacher distillation from diverse frontier models while admitting multiple valid implementation strategies.",
"arxiv_id": "2601.09165",
"authors": [
"Aaron R. Flouro",
"Shawn P. Chadwick"
],
"categories": [
"cs.LG"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Multi-Teacher Ensemble Distillation: A Mathematical Framework for Probability-Domain Knowledge Aggregation",
"url": "https://arxiv.org/abs/2601.09165",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "b28b29e9-96f3-45c3-a37b-740b297e6233",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}