dorsal/arxiv
View SchemaChinese Labor Law Large Language Model Benchmark
| Authors | Zixun Lan, Maochun Xu, Yifan Ren, Rui Wu, Jianghui Zhou, Xueyang Cheng, Jianan Ding Ding, Xinheng Wang, Mingmin Chi, Fei Ma |
|---|---|
| Categories | |
| ArXiv ID | 2601.09972vv1 |
| URL | https://arxiv.org/abs/2601.09972 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains that require precise legal knowledge, complex reasoning, and contextual sensitivity. To address these limitations, we present LabourLawLLM, a legal large language model tailored to Chinese labor law. We also introduce LabourLawBench, a comprehensive benchmark covering diverse labor-law tasks, including legal provision citation, knowledge-based question answering, case classification, compensation computation, named entity recognition, and legal case analysis. Our evaluation framework combines objective metrics (e.g., ROUGE-L, accuracy, F1, and soft-F1) with subjective assessment based on GPT-4 scoring. Experiments show that LabourLawLLM consistently outperforms general-purpose and existing legal-specific LLMs across task categories. Beyond labor law, our methodology provides a scalable approach for building specialized LLMs in other legal subfields, improving accuracy, reliability, and societal value of legal AI applications.
{
"annotation_id": "ab274625-9b9d-42a5-9e1c-1cb120e66798",
"date_created": "2026-02-17T05:53:24.394000Z",
"date_modified": "2026-02-17T05:53:24.394000Z",
"file_hash": "bcda936134b25363376c382a2b23888f3fb89fffb39944aacba971647a6bf833",
"private": false,
"record": {
"abstract": "Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains that require precise legal knowledge, complex reasoning, and contextual sensitivity. To address these limitations, we present LabourLawLLM, a legal large language model tailored to Chinese labor law. We also introduce LabourLawBench, a comprehensive benchmark covering diverse labor-law tasks, including legal provision citation, knowledge-based question answering, case classification, compensation computation, named entity recognition, and legal case analysis. Our evaluation framework combines objective metrics (e.g., ROUGE-L, accuracy, F1, and soft-F1) with subjective assessment based on GPT-4 scoring. Experiments show that LabourLawLLM consistently outperforms general-purpose and existing legal-specific LLMs across task categories. Beyond labor law, our methodology provides a scalable approach for building specialized LLMs in other legal subfields, improving accuracy, reliability, and societal value of legal AI applications.",
"arxiv_id": "2601.09972",
"authors": [
"Zixun Lan",
"Maochun Xu",
"Yifan Ren",
"Rui Wu",
"Jianghui Zhou",
"Xueyang Cheng",
"Jianan Ding Ding",
"Xinheng Wang",
"Mingmin Chi",
"Fei Ma"
],
"categories": [
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Chinese Labor Law Large Language Model Benchmark",
"url": "https://arxiv.org/abs/2601.09972",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "a3f39e38-74b2-47b1-a843-3c51fba2837c",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}