dorsal/arxiv
View SchemaBeyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation
| Authors | Yanzhi Tian, Cunxiang Wang, Zeming Liu, Heyan Huang, Wenbo Yu, Dawei Song, Jie Tang, Yuhang Guo |
|---|---|
| Categories | |
| ArXiv ID | 2601.07338vv1 |
| URL | https://arxiv.org/abs/2601.07338 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Large Language Models (LLMs) have significantly advanced Machine Translation (MT), applying them to linguistically complex domains-such as Social Network Services, literature etc. In these scenarios, translations often require handling non-literal expressions, leading to the inaccuracy of MT metrics. To systematically investigate the reliability of MT metrics, we first curate a meta-evaluation dataset focused on non-literal translations, namely MENT. MENT encompasses four non-literal translation domains and features source sentences paired with translations from diverse MT systems, with 7,530 human-annotated scores on translation quality. Experimental results reveal the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge, particularly the knowledge cutoff and score inconsistency problem. To mitigate these limitations, we propose RATE, a novel agentic translation evaluation framework, centered by a reflective Core Agent that dynamically invokes specialized sub-agents. Experimental results indicate the efficacy of RATE, achieving an improvement of at least 3.2 meta score compared with current metrics. Further experiments demonstrate the robustness of RATE to general-domain MT evaluation. Code and dataset are available at: https://github.com/BITHLP/RATE.
{
"annotation_id": "d4062776-cfb6-41d5-a82e-925c10112ac4",
"date_created": "2026-02-17T05:53:12.658000Z",
"date_modified": "2026-02-17T05:53:12.658000Z",
"file_hash": "cfcf53665222733da40a4d8c52160ae381a8be5b8f89644e550329fd3ae3a0e2",
"private": false,
"record": {
"abstract": "Large Language Models (LLMs) have significantly advanced Machine Translation (MT), applying them to linguistically complex domains-such as Social Network Services, literature etc. In these scenarios, translations often require handling non-literal expressions, leading to the inaccuracy of MT metrics. To systematically investigate the reliability of MT metrics, we first curate a meta-evaluation dataset focused on non-literal translations, namely MENT. MENT encompasses four non-literal translation domains and features source sentences paired with translations from diverse MT systems, with 7,530 human-annotated scores on translation quality. Experimental results reveal the inaccuracies of traditional MT metrics and the limitations of LLM-as-a-Judge, particularly the knowledge cutoff and score inconsistency problem. To mitigate these limitations, we propose RATE, a novel agentic translation evaluation framework, centered by a reflective Core Agent that dynamically invokes specialized sub-agents. Experimental results indicate the efficacy of RATE, achieving an improvement of at least 3.2 meta score compared with current metrics. Further experiments demonstrate the robustness of RATE to general-domain MT evaluation. Code and dataset are available at: https://github.com/BITHLP/RATE.",
"arxiv_id": "2601.07338",
"authors": [
"Yanzhi Tian",
"Cunxiang Wang",
"Zeming Liu",
"Heyan Huang",
"Wenbo Yu",
"Dawei Song",
"Jie Tang",
"Yuhang Guo"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Beyond Literal Mapping: Benchmarking and Improving Non-Literal Translation Evaluation",
"url": "https://arxiv.org/abs/2601.07338",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "48ca317b-b5b3-445d-8b98-3750b70acbdd",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}