dorsal/arxiv
View SchemaEfficient Multilingual Dialogue Processing via Translation Pipelines and Distilled Language Models
| Authors | Santiago Martínez Novoa, Nicolás Rozo Fajardo, Diego Alejandro González Vargas, Nicolás Bedoya Figueroa |
|---|---|
| Categories | |
| ArXiv ID | 2601.09059vv1 |
| URL | https://arxiv.org/abs/2601.09059 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
This paper presents team Kl33n3x's multilingual dialogue summarization and question answering system developed for the NLPAI4Health 2025 shared task. The approach employs a three-stage pipeline: forward translation from Indic languages to English, multitask text generation using a 2.55B parameter distilled language model, and reverse translation back to source languages. By leveraging knowledge distillation techniques, this work demonstrates that compact models can achieve highly competitive performance across nine languages. The system achieved strong win rates across the competition's tasks, with particularly robust performance on Marathi (86.7% QnA), Tamil (86.7% QnA), and Hindi (80.0% QnA), demonstrating the effectiveness of translation-based approaches for low-resource language processing without task-specific fine-tuning.
{
"annotation_id": "db33027f-10d1-4f4f-9ca3-b39ed4f006d5",
"date_created": "2026-02-17T05:53:19.348000Z",
"date_modified": "2026-02-17T05:53:19.348000Z",
"file_hash": "6ce778476e5c54f19d07b7dae69b70d62c45c7c65bf1c50c18bb4cd5edfd39ae",
"private": false,
"record": {
"abstract": "This paper presents team Kl33n3x\u0027s multilingual dialogue summarization and question answering system developed for the NLPAI4Health 2025 shared task. The approach employs a three-stage pipeline: forward translation from Indic languages to English, multitask text generation using a 2.55B parameter distilled language model, and reverse translation back to source languages. By leveraging knowledge distillation techniques, this work demonstrates that compact models can achieve highly competitive performance across nine languages. The system achieved strong win rates across the competition\u0027s tasks, with particularly robust performance on Marathi (86.7% QnA), Tamil (86.7% QnA), and Hindi (80.0% QnA), demonstrating the effectiveness of translation-based approaches for low-resource language processing without task-specific fine-tuning.",
"arxiv_id": "2601.09059",
"authors": [
"Santiago Mart\u00ednez Novoa",
"Nicol\u00e1s Rozo Fajardo",
"Diego Alejandro Gonz\u00e1lez Vargas",
"Nicol\u00e1s Bedoya Figueroa"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Efficient Multilingual Dialogue Processing via Translation Pipelines and Distilled Language Models",
"url": "https://arxiv.org/abs/2601.09059",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "9f78009a-8697-4311-9fc5-87ab31328ae3",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}