dorsal/arxiv
View SchemaPERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models
| Authors | Chengbing Wang, Wuqiang Zheng, Yang Zhang, Fengbin Zhu, Junyi Cheng, Yi Xie, Wenjie Wang, Fuli Feng |
|---|---|
| Categories | |
| ArXiv ID | 2601.10532vv1 |
| URL | https://arxiv.org/abs/2601.10532 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Large Language Models (LLMs) are increasingly deployed in human-centric applications, yet they often fail to provide substantive emotional support. While Reinforcement Learning (RL) has been utilized to enhance empathy of LLMs, existing reward models typically evaluate empathy from a single perspective, overlooking the inherently bidirectional interaction nature of empathy between the supporter and seeker as defined by Empathy Cycle theory. To address this limitation, we propose Psychology-grounded Empathetic Reward Modeling (PERM). PERM operationalizes empathy evaluation through a bidirectional decomposition: 1) Supporter perspective, assessing internal resonation and communicative expression; 2) Seeker perspective, evaluating emotional reception. Additionally, it incorporates a bystander perspective to monitor overall interaction quality. Extensive experiments on a widely-used emotional intelligence benchmark and an industrial daily conversation dataset demonstrate that PERM outperforms state-of-the-art baselines by over 10\%. Furthermore, a blinded user study reveals a 70\% preference for our approach, highlighting its efficacy in generating more empathetic responses. Our code, dataset, and models are available at https://github.com/ZhengWwwq/PERM.
{
"annotation_id": "f7f5b132-dc9a-4d28-8e9a-2fbe28f6026e",
"date_created": "2026-02-17T05:53:24.058000Z",
"date_modified": "2026-02-17T05:53:24.058000Z",
"file_hash": "7ead164ee9fc4e3d3f3d69631024b45b2e6af6aea7cabcb6cbe6a826e1410104",
"private": false,
"record": {
"abstract": "Large Language Models (LLMs) are increasingly deployed in human-centric applications, yet they often fail to provide substantive emotional support. While Reinforcement Learning (RL) has been utilized to enhance empathy of LLMs, existing reward models typically evaluate empathy from a single perspective, overlooking the inherently bidirectional interaction nature of empathy between the supporter and seeker as defined by Empathy Cycle theory. To address this limitation, we propose Psychology-grounded Empathetic Reward Modeling (PERM). PERM operationalizes empathy evaluation through a bidirectional decomposition: 1) Supporter perspective, assessing internal resonation and communicative expression; 2) Seeker perspective, evaluating emotional reception. Additionally, it incorporates a bystander perspective to monitor overall interaction quality. Extensive experiments on a widely-used emotional intelligence benchmark and an industrial daily conversation dataset demonstrate that PERM outperforms state-of-the-art baselines by over 10\\%. Furthermore, a blinded user study reveals a 70\\% preference for our approach, highlighting its efficacy in generating more empathetic responses. Our code, dataset, and models are available at https://github.com/ZhengWwwq/PERM.",
"arxiv_id": "2601.10532",
"authors": [
"Chengbing Wang",
"Wuqiang Zheng",
"Yang Zhang",
"Fengbin Zhu",
"Junyi Cheng",
"Yi Xie",
"Wenjie Wang",
"Fuli Feng"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "PERM: Psychology-grounded Empathetic Reward Modeling for Large Language Models",
"url": "https://arxiv.org/abs/2601.10532",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "e4d09fd6-6255-4224-bb8e-8c9e6e71974b",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}