dorsal/arxiv
View SchemaAssessing and Improving Punctuation Robustness in English-Marathi Machine Translation
| Authors | Kaustubh Shivshankar Shejole, Sourabh Deoghare, Pushpak Bhattacharyya |
|---|---|
| Categories | |
| ArXiv ID | 2601.09725vv2 |
| URL | https://arxiv.org/abs/2601.09725 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Punctuation plays a critical role in resolving semantic and structural ambiguity in written language. Machine Translation (MT) systems are now widely applied across diverse domains and languages, including many low-resource settings. In this work, we focus on Marathi, a low- to middle-resource language. We introduce Vir\=am, the first diagnostic benchmark for assessing punctuation robustness in English-to-Marathi machine translation, consisting of 54 manually curated, punctuation-ambiguous instances. We evaluate two primary strategies for enhancing reliability: a pipeline-based restore-then-translate approach and direct fine-tuned on punctuation-varied data. Our results demonstrate that specialized fine-tuned models and pipeline systems significantly improve translation quality over standard baselines on the Vir\=am benchmark. Qualitative analysis reveals that the original model may result in wrong translations leading to wrong interpretations, while fine-tuned models significantly improve overall reliability. Furthermore, we find that current Large Language Models (LLMs) lag behind these task-specific approaches in preserving meaning for punctuation-ambiguous text, thus necessitating further research in this area.
{
"annotation_id": "d6b1b4c3-59d8-4695-b07c-624ec78912c7",
"date_created": "2026-02-17T05:53:24.376000Z",
"date_modified": "2026-02-17T05:53:24.376000Z",
"file_hash": "58622e33c62179873854ffee8998b233c6f620ebe0f4b7d83bb412f314ad5d68",
"private": false,
"record": {
"abstract": "Punctuation plays a critical role in resolving semantic and structural ambiguity in written language. Machine Translation (MT) systems are now widely applied across diverse domains and languages, including many low-resource settings. In this work, we focus on Marathi, a low- to middle-resource language. We introduce Vir\\=am, the first diagnostic benchmark for assessing punctuation robustness in English-to-Marathi machine translation, consisting of 54 manually curated, punctuation-ambiguous instances. We evaluate two primary strategies for enhancing reliability: a pipeline-based restore-then-translate approach and direct fine-tuned on punctuation-varied data. Our results demonstrate that specialized fine-tuned models and pipeline systems significantly improve translation quality over standard baselines on the Vir\\=am benchmark. Qualitative analysis reveals that the original model may result in wrong translations leading to wrong interpretations, while fine-tuned models significantly improve overall reliability. Furthermore, we find that current Large Language Models (LLMs) lag behind these task-specific approaches in preserving meaning for punctuation-ambiguous text, thus necessitating further research in this area.",
"arxiv_id": "2601.09725",
"authors": [
"Kaustubh Shivshankar Shejole",
"Sourabh Deoghare",
"Pushpak Bhattacharyya"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Assessing and Improving Punctuation Robustness in English-Marathi Machine Translation",
"url": "https://arxiv.org/abs/2601.09725",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "d2d48594-4aa6-4024-bf47-1f9968ac91c2",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}