dorsal/arxiv
View SchemaAssessing and Improving Punctuation Robustness in English-Marathi Machine Translation
| Authors | Kaustubh Shivshankar Shejole, Sourabh Deoghare, Pushpak Bhattacharyya |
|---|---|
| Categories | |
| ArXiv ID | 2601.09725vv1 |
| URL | https://arxiv.org/abs/2601.09725 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Punctuation plays a critical role in resolving semantic and structural ambiguity in written language. Machine Translation (MT) systems are now widely applied across diverse domains and languages, including many low-resource settings. In this work, we focus on Marathi, a low- to middle-resource language. We introduce Vir\=am, the first diagnostic benchmark for assessing punctuation robustness in English-to-Marathi machine translation, consisting of 54 manually curated, punctuation-ambiguous instances. We evaluate two primary strategies for enhancing reliability: a pipeline-based restore-then-translate approach and direct fine-tuned on punctuation-varied data. Our results demonstrate that specialized fine-tuned models and pipeline systems significantly improve translation quality over standard baselines on the Vir\=am benchmark. Qualitative analysis reveals that the original model may result in wrong translations leading to wrong interpretations, while fine-tuned models significantly improve overall reliability. Furthermore, we find that current Large Language Models (LLMs) lag behind these task-specific approaches in preserving meaning for punctuation-ambiguous text, thus necessitating further research in this area.
{
"annotation_id": "84b6c97d-7a52-4ac6-bf38-9c2db5ea4f40",
"date_created": "2026-02-17T05:53:23.006000Z",
"date_modified": "2026-02-17T05:53:23.006000Z",
"file_hash": "20afbf2fb5dd6287d2dd97264589bb658b940f5afb11a370fe2631d7be068851",
"private": false,
"record": {
"abstract": "Punctuation plays a critical role in resolving semantic and structural ambiguity in written language. Machine Translation (MT) systems are now widely applied across diverse domains and languages, including many low-resource settings. In this work, we focus on Marathi, a low- to middle-resource language. We introduce Vir\\=am, the first diagnostic benchmark for assessing punctuation robustness in English-to-Marathi machine translation, consisting of 54 manually curated, punctuation-ambiguous instances. We evaluate two primary strategies for enhancing reliability: a pipeline-based restore-then-translate approach and direct fine-tuned on punctuation-varied data. Our results demonstrate that specialized fine-tuned models and pipeline systems significantly improve translation quality over standard baselines on the Vir\\=am benchmark. Qualitative analysis reveals that the original model may result in wrong translations leading to wrong interpretations, while fine-tuned models significantly improve overall reliability. Furthermore, we find that current Large Language Models (LLMs) lag behind these task-specific approaches in preserving meaning for punctuation-ambiguous text, thus necessitating further research in this area.",
"arxiv_id": "2601.09725",
"authors": [
"Kaustubh Shivshankar Shejole",
"Sourabh Deoghare",
"Pushpak Bhattacharyya"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Assessing and Improving Punctuation Robustness in English-Marathi Machine Translation",
"url": "https://arxiv.org/abs/2601.09725",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "fdfccd1f-c01a-4404-8885-b64108c03ab4",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}