dorsal/arxiv
View SchemaKnowing But Not Doing: Convergent Morality and Divergent Action in LLMs
| Authors | Jen-tse Huang, Jiantong Qin, Xueli Qiu, Sharon Levy, Michelle R. Kaufman, Mark Dredze |
|---|---|
| Categories | |
| ArXiv ID | 2601.07972vv1 |
| URL | https://arxiv.org/abs/2601.07972 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in real-world decision contexts remains under-explored. We present ValAct-15k, a dataset of 3,000 advice-seeking scenarios derived from Reddit, designed to elicit ten values defined by Schwartz Theory of Basic Human Values. Using both the scenario-based questions and the traditional value questionnaire, we evaluate ten frontier LLMs (five from U.S. companies, five from Chinese ones) and human participants ($n = 55$). We find near-perfect cross-model consistency in scenario-based decisions (Pearson $r \approx 1.0$), contrasting sharply with the broad variability observed among humans ($r \in [-0.79, 0.98]$). Yet, both humans and LLMs show weak correspondence between self-reported and enacted values ($r = 0.4, 0.3$), revealing a systematic knowledge-action gap. When instructed to "hold" a specific value, LLMs' performance declines up to $6.6%$ compared to merely selecting the value, indicating a role-play aversion. These findings suggest that while alignment training yields normative value convergence, it does not eliminate the human-like incoherence between knowing and acting upon values.
{
"annotation_id": "7e51d1c3-2ff2-4b37-89db-86a40b412267",
"date_created": "2026-02-17T05:53:15.044000Z",
"date_modified": "2026-02-17T05:53:15.044000Z",
"file_hash": "b0b56907834cd4ede7556395ffd317e3ae46d92b086108f5a1dcba41edb2ad3c",
"private": false,
"record": {
"abstract": "Value alignment is central to the development of safe and socially compatible artificial intelligence. However, how Large Language Models (LLMs) represent and enact human values in real-world decision contexts remains under-explored. We present ValAct-15k, a dataset of 3,000 advice-seeking scenarios derived from Reddit, designed to elicit ten values defined by Schwartz Theory of Basic Human Values. Using both the scenario-based questions and the traditional value questionnaire, we evaluate ten frontier LLMs (five from U.S. companies, five from Chinese ones) and human participants ($n = 55$). We find near-perfect cross-model consistency in scenario-based decisions (Pearson $r \\approx 1.0$), contrasting sharply with the broad variability observed among humans ($r \\in [-0.79, 0.98]$). Yet, both humans and LLMs show weak correspondence between self-reported and enacted values ($r = 0.4, 0.3$), revealing a systematic knowledge-action gap. When instructed to \"hold\" a specific value, LLMs\u0027 performance declines up to $6.6%$ compared to merely selecting the value, indicating a role-play aversion. These findings suggest that while alignment training yields normative value convergence, it does not eliminate the human-like incoherence between knowing and acting upon values.",
"arxiv_id": "2601.07972",
"authors": [
"Jen-tse Huang",
"Jiantong Qin",
"Xueli Qiu",
"Sharon Levy",
"Michelle R. Kaufman",
"Mark Dredze"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Knowing But Not Doing: Convergent Morality and Divergent Action in LLMs",
"url": "https://arxiv.org/abs/2601.07972",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "5b4fce8e-2781-4c4c-9379-fb77b40779cb",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}