dorsal/arxiv
View SchemaWhere Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
| Authors | Minh Vu Pham, Hsuvas Borkakoty, Yufang Hou |
|---|---|
| Categories | |
| ArXiv ID | 2601.09445vv1 |
| URL | https://arxiv.org/abs/2601.09445 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
In language models (LMs), intra-memory knowledge conflict largely arises when inconsistent information about the same event is encoded within the model's parametric knowledge. While prior work has primarily focused on resolving conflicts between a model's internal knowledge and external resources through approaches such as fine-tuning or knowledge editing, the problem of localizing conflicts that originate during pre-training within the model's internal representations remain unexplored. In this work, we design a framework based on mechanistic interpretability methods to identify where and how conflicting knowledge from the pre-training data is encoded within LMs. Our findings contribute to a growing body of evidence that specific internal components of a language model are responsible for encoding conflicting knowledge from pre-training, and we demonstrate how mechanistic interpretability methods can be leveraged to causally intervene in and control conflicting knowledge at inference time.
{
"annotation_id": "149b4e8c-da63-4387-b7c4-8e5d8d80bda4",
"date_created": "2026-02-17T05:53:20.118000Z",
"date_modified": "2026-02-17T05:53:20.118000Z",
"file_hash": "b929efcfb33742ff719892d9d7c78d7890e3df5f4c99073264c9b7f709465f50",
"private": false,
"record": {
"abstract": "In language models (LMs), intra-memory knowledge conflict largely arises when inconsistent information about the same event is encoded within the model\u0027s parametric knowledge. While prior work has primarily focused on resolving conflicts between a model\u0027s internal knowledge and external resources through approaches such as fine-tuning or knowledge editing, the problem of localizing conflicts that originate during pre-training within the model\u0027s internal representations remain unexplored. In this work, we design a framework based on mechanistic interpretability methods to identify where and how conflicting knowledge from the pre-training data is encoded within LMs. Our findings contribute to a growing body of evidence that specific internal components of a language model are responsible for encoding conflicting knowledge from pre-training, and we demonstrate how mechanistic interpretability methods can be leveraged to causally intervene in and control conflicting knowledge at inference time.",
"arxiv_id": "2601.09445",
"authors": [
"Minh Vu Pham",
"Hsuvas Borkakoty",
"Yufang Hou"
],
"categories": [
"cs.CL",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models",
"url": "https://arxiv.org/abs/2601.09445",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "43bf6ddc-4ffe-46c2-922b-e82fd0c7fa91",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}