dorsal/arxiv
View SchemaReasoning Models Will Blatantly Lie About Their Reasoning
| Authors | William Walden |
|---|---|
| Categories | |
| ArXiv ID | 2601.07663vv2 |
| URL | https://arxiv.org/abs/2601.07663 |
| License | http://creativecommons.org/licenses/by-sa/4.0/ |
Abstract
It has been shown that Large Reasoning Models (LRMs) may not *say what they think*: they do not always volunteer information about how certain parts of the input influence their reasoning. But it is one thing for a model to *omit* such information and another, worse thing to *lie* about it. Here, we extend the work of Chen et al. (2025) to show that LRMs will do just this: they will flatly deny relying on hints provided in the prompt in answering multiple choice questions -- even when directly asked to reflect on unusual (i.e. hinted) prompt content, even when allowed to use hints, and even though experiments *show* them to be using the hints. Our results thus have discouraging implications for CoT monitoring and interpretability.
{
"annotation_id": "d34dc5a9-3312-487e-ae57-96ebbdfd74aa",
"date_created": "2026-02-17T05:53:12.489000Z",
"date_modified": "2026-02-17T05:53:12.489000Z",
"file_hash": "86d6d529e11612d635f77f5eb7c2d4fcf7f1a395d529204bc7a9e2a609d02c3d",
"private": false,
"record": {
"abstract": "It has been shown that Large Reasoning Models (LRMs) may not *say what they think*: they do not always volunteer information about how certain parts of the input influence their reasoning. But it is one thing for a model to *omit* such information and another, worse thing to *lie* about it. Here, we extend the work of Chen et al. (2025) to show that LRMs will do just this: they will flatly deny relying on hints provided in the prompt in answering multiple choice questions -- even when directly asked to reflect on unusual (i.e. hinted) prompt content, even when allowed to use hints, and even though experiments *show* them to be using the hints. Our results thus have discouraging implications for CoT monitoring and interpretability.",
"arxiv_id": "2601.07663",
"authors": [
"William Walden"
],
"categories": [
"cs.AI",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by-sa/4.0/",
"title": "Reasoning Models Will Blatantly Lie About Their Reasoning",
"url": "https://arxiv.org/abs/2601.07663",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "7c9ea5e4-acd7-4d56-ac69-cb682b82d905",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}