dorsal/arxiv
View SchemaReasoning Models Will Blatantly Lie About Their Reasoning
| Authors | William Walden |
|---|---|
| Categories | |
| ArXiv ID | 2601.07663vv1 |
| URL | https://arxiv.org/abs/2601.07663 |
| License | http://creativecommons.org/licenses/by-sa/4.0/ |
Abstract
It has been shown that Large Reasoning Models (LRMs) may not *say what they think*: they do not always volunteer information about how certain parts of the input influence their reasoning. But it is one thing for a model to *omit* such information and another, worse thing to *lie* about it. Here, we extend the work of Chen et al. (2025) to show that LRMs will do just this: they will flatly deny relying on hints provided in the prompt in answering multiple choice questions -- even when directly asked to reflect on unusual (i.e. hinted) prompt content, even when allowed to use hints, and even though experiments *show* them to be using the hints. Our results thus have discouraging implications for CoT monitoring and interpretability.
{
"annotation_id": "6e8e4e2e-004a-424c-b54e-e91a9b1b09d8",
"date_created": "2026-02-17T05:53:12.485000Z",
"date_modified": "2026-02-17T05:53:12.485000Z",
"file_hash": "f4aabfad91ea1415dffea5d5bb2dbfeef18d054a35618e28677518e140610f0f",
"private": false,
"record": {
"abstract": "It has been shown that Large Reasoning Models (LRMs) may not *say what they think*: they do not always volunteer information about how certain parts of the input influence their reasoning. But it is one thing for a model to *omit* such information and another, worse thing to *lie* about it. Here, we extend the work of Chen et al. (2025) to show that LRMs will do just this: they will flatly deny relying on hints provided in the prompt in answering multiple choice questions -- even when directly asked to reflect on unusual (i.e. hinted) prompt content, even when allowed to use hints, and even though experiments *show* them to be using the hints. Our results thus have discouraging implications for CoT monitoring and interpretability.",
"arxiv_id": "2601.07663",
"authors": [
"William Walden"
],
"categories": [
"cs.AI",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by-sa/4.0/",
"title": "Reasoning Models Will Blatantly Lie About Their Reasoning",
"url": "https://arxiv.org/abs/2601.07663",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "1ef32dc3-79b6-4fa9-b96c-126b5fadd272",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}