dorsal/arxiv
View SchemaMulti-Level Embedding Conformer Framework for Bengali Automatic Speech Recognition
| Authors | Md. Nazmus Sakib, Golam Mahmud, Md. Maruf Bangabashi, Umme Ara Mahinur Istia, Md. Jahidul Islam, Partha Sarker, Afra Yeamini Prity |
|---|---|
| Categories | |
| ArXiv ID | 2601.09710vv1 |
| URL | https://arxiv.org/abs/2601.09710 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Bengali, spoken by over 300 million people, is a morphologically rich and lowresource language, posing challenges for automatic speech recognition (ASR). This research presents an end-to-end framework for Bengali ASR, building on a Conformer-CTC backbone with a multi-level embedding fusion mechanism that incorporates phoneme, syllable, and wordpiece representations. By enriching acoustic features with these linguistic embeddings, the model captures fine-grained phonetic cues and higher-level contextual patterns. The architecture employs early and late Conformer stages, with preprocessing steps including silence trimming, resampling, Log-Mel spectrogram extraction, and SpecAugment augmentation. The experimental results demonstrate the strong potential of the model, achieving a word error rate (WER) of 10.01% and a character error rate (CER) of 5.03%. These results demonstrate the effectiveness of combining multi-granular linguistic information with acoustic modeling, providing a scalable approach for low-resource ASR development.
{
"annotation_id": "7883a00c-8511-45e6-8e59-bfc32f269cfa",
"date_created": "2026-02-17T05:53:20.175000Z",
"date_modified": "2026-02-17T05:53:20.175000Z",
"file_hash": "858e90754f85f3b2b5cbf02eeb3e8dda9ed71a325e18b5915b432123b1c508f2",
"private": false,
"record": {
"abstract": "Bengali, spoken by over 300 million people, is a morphologically rich and lowresource language, posing challenges for automatic speech recognition (ASR). This research presents an end-to-end framework for Bengali ASR, building on a Conformer-CTC backbone with a multi-level embedding fusion mechanism that incorporates phoneme, syllable, and wordpiece representations. By enriching acoustic features with these linguistic embeddings, the model captures fine-grained phonetic cues and higher-level contextual patterns. The architecture employs early and late Conformer stages, with preprocessing steps including silence trimming, resampling, Log-Mel spectrogram extraction, and SpecAugment augmentation. The experimental results demonstrate the strong potential of the model, achieving a word error rate (WER) of 10.01% and a character error rate (CER) of 5.03%. These results demonstrate the effectiveness of combining multi-granular linguistic information with acoustic modeling, providing a scalable approach for low-resource ASR development.",
"arxiv_id": "2601.09710",
"authors": [
"Md. Nazmus Sakib",
"Golam Mahmud",
"Md. Maruf Bangabashi",
"Umme Ara Mahinur Istia",
"Md. Jahidul Islam",
"Partha Sarker",
"Afra Yeamini Prity"
],
"categories": [
"eess.AS",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Multi-Level Embedding Conformer Framework for Bengali Automatic Speech Recognition",
"url": "https://arxiv.org/abs/2601.09710",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "aafb0b42-5d85-46bb-850d-8adfd5958d51",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}