dorsal/arxiv
View SchemaSITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages
| Authors | Tianyi Xu, Xuan Ouyang, Binwei Yao, Shoua Xiong, Sara Misurelli, Maichou Lor, Junjie Hu |
|---|---|
| Categories | |
| ArXiv ID | 2601.09050vv1 |
| URL | https://arxiv.org/abs/2601.09050 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Tonal low-resource languages are widely spoken yet remain underserved by modern speech technology. A key challenge is learning representations that are robust to nuisance variation such as gender while remaining tone-aware for different lexical meanings. To address this, we propose SITA, a lightweight adaptation recipe that enforces Speaker-Invariance and Tone-Awareness for pretrained wav2vec-style encoders. SITA uses staged multi-objective training: (i) a cross-gender contrastive objective encourages lexical consistency across speakers, while a tone-repulsive loss prevents tone collapse by explicitly separating same-word different-tone realizations; and (ii) an auxiliary Connectionist Temporal Classification (CTC)-based ASR objective with distillation stabilizes recognition-relevant structure. We evaluate primarily on Hmong, a highly tonal and severely under-resourced language where off-the-shelf multilingual encoders fail to represent tone effectively. On a curated Hmong word corpus, SITA improves cross-gender lexical retrieval accuracy, while maintaining usable ASR accuracy relative to an ASR-adapted XLS-R teacher. We further observe similar gains when transferring the same recipe to Mandarin, suggesting SITA is a general, plug-in approach for adapting multilingual speech encoders to tonal languages.
{
"annotation_id": "fa3a6dfe-0eb1-4419-bda9-ee381b3539be",
"date_created": "2026-02-17T05:53:20.469000Z",
"date_modified": "2026-02-17T05:53:20.469000Z",
"file_hash": "fa25ca8aa9b2071031322649dac29966d13f4b14e46940589c099df6ed51b2f5",
"private": false,
"record": {
"abstract": "Tonal low-resource languages are widely spoken yet remain underserved by modern speech technology. A key challenge is learning representations that are robust to nuisance variation such as gender while remaining tone-aware for different lexical meanings. To address this, we propose SITA, a lightweight adaptation recipe that enforces Speaker-Invariance and Tone-Awareness for pretrained wav2vec-style encoders. SITA uses staged multi-objective training: (i) a cross-gender contrastive objective encourages lexical consistency across speakers, while a tone-repulsive loss prevents tone collapse by explicitly separating same-word different-tone realizations; and (ii) an auxiliary Connectionist Temporal Classification (CTC)-based ASR objective with distillation stabilizes recognition-relevant structure. We evaluate primarily on Hmong, a highly tonal and severely under-resourced language where off-the-shelf multilingual encoders fail to represent tone effectively. On a curated Hmong word corpus, SITA improves cross-gender lexical retrieval accuracy, while maintaining usable ASR accuracy relative to an ASR-adapted XLS-R teacher. We further observe similar gains when transferring the same recipe to Mandarin, suggesting SITA is a general, plug-in approach for adapting multilingual speech encoders to tonal languages.",
"arxiv_id": "2601.09050",
"authors": [
"Tianyi Xu",
"Xuan Ouyang",
"Binwei Yao",
"Shoua Xiong",
"Sara Misurelli",
"Maichou Lor",
"Junjie Hu"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "SITA: Learning Speaker-Invariant and Tone-Aware Speech Representations for Low-Resource Tonal Languages",
"url": "https://arxiv.org/abs/2601.09050",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "678eb533-b454-4a6d-bf54-d76cfe35cf4c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}