dorsal/arxiv
View SchemaTowards Native Intelligence: 6G-LLM Trained with Reinforcement Learning from NDT Feedback
| Authors | Zhuoran Xiao, Tao Tao, Chenhui Ye, Yunbo Hu, Yijia Feng, Tianyu Jiao, Liyu Cai |
|---|---|
| Categories | |
| ArXiv ID | 2601.09992vv1 |
| URL | https://arxiv.org/abs/2601.09992 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Owing to its comprehensive understanding of upper-layer application requirements and the capabilities of practical communication systems, the 6G-LLM (6G domain large language model) offers a promising pathway toward realizing network native intelligence. Serving as the system orchestrator, the 6G-LLM drives a paradigm shift that fundamentally departs from existing rule-based approaches, which primarily rely on modular, experience-driven optimization. By contrast, the 6G-LLM substantially enhances network flexibility and adaptability. Nevertheless, current efforts to construct 6G-LLMs are constrained by their reliance on large-scale, meticulously curated, human-authored corpora, which are impractical to obtain in real-world scenarios. Moreover, purely offline-trained models lack the capacity for continual self-improvement, limiting their ability to adapt to the highly dynamic requirements of wireless communication environments. To overcome these limitations, we propose a novel training paradigm termed RLDTF (Reinforcement Learning from Digital Twin Feedback) for 6G-LLMs. This framework leverages network digital twins to generate reward signals based on orchestration outcomes, while employing reinforcement learning to guide the model toward optimal decision-making dynamically. Furthermore, we introduce a weighted token mechanism to improve output accuracy. Comprehensive experimental results demonstrate that our proposed framework significantly outperforms state-of-the-art baselines in orchestration accuracy and solution optimality.
{
"annotation_id": "e8591fdb-4252-40c6-bd14-c9c68e97fd97",
"date_created": "2026-02-17T05:53:24.223000Z",
"date_modified": "2026-02-17T05:53:24.223000Z",
"file_hash": "4518f668a3054b4b6893fe552c4890ab1a8ae2d398811472cbf9eebe76f18246",
"private": false,
"record": {
"abstract": "Owing to its comprehensive understanding of upper-layer application requirements and the capabilities of practical communication systems, the 6G-LLM (6G domain large language model) offers a promising pathway toward realizing network native intelligence. Serving as the system orchestrator, the 6G-LLM drives a paradigm shift that fundamentally departs from existing rule-based approaches, which primarily rely on modular, experience-driven optimization. By contrast, the 6G-LLM substantially enhances network flexibility and adaptability. Nevertheless, current efforts to construct 6G-LLMs are constrained by their reliance on large-scale, meticulously curated, human-authored corpora, which are impractical to obtain in real-world scenarios. Moreover, purely offline-trained models lack the capacity for continual self-improvement, limiting their ability to adapt to the highly dynamic requirements of wireless communication environments. To overcome these limitations, we propose a novel training paradigm termed RLDTF (Reinforcement Learning from Digital Twin Feedback) for 6G-LLMs. This framework leverages network digital twins to generate reward signals based on orchestration outcomes, while employing reinforcement learning to guide the model toward optimal decision-making dynamically. Furthermore, we introduce a weighted token mechanism to improve output accuracy. Comprehensive experimental results demonstrate that our proposed framework significantly outperforms state-of-the-art baselines in orchestration accuracy and solution optimality.",
"arxiv_id": "2601.09992",
"authors": [
"Zhuoran Xiao",
"Tao Tao",
"Chenhui Ye",
"Yunbo Hu",
"Yijia Feng",
"Tianyu Jiao",
"Liyu Cai"
],
"categories": [
"eess.SP"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Towards Native Intelligence: 6G-LLM Trained with Reinforcement Learning from NDT Feedback",
"url": "https://arxiv.org/abs/2601.09992",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "cf8dcd03-3f4f-4afe-8e53-05c8759a7b13",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}