dorsal/arxiv
View SchemaToward Understanding Unlearning Difficulty: A Mechanistic Perspective and Circuit-Guided Difficulty Metric
| Authors | Jiali Cheng, Ziheng Chen, Chirag Agarwal, Hadi Amiri |
|---|---|
| Categories | |
| ArXiv ID | 2601.09624vv1 |
| URL | https://arxiv.org/abs/2601.09624 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably across individual samples: some are reliably erased, while others persist despite the same procedure. We argue that this disparity is not only a data-side phenomenon, but also reflects model-internal mechanisms that encode and protect memorized information. We study this problem from a mechanistic perspective based on model circuits--structured interaction pathways that govern how predictions are formed. We propose Circuit-guided Unlearning Difficulty (CUD), a {\em pre-unlearning} metric that assigns each sample a continuous difficulty score using circuit-level signals. Extensive experiments demonstrate that CUD reliably separates intrinsically easy and hard samples, and remains stable across unlearning methods. We identify key circuit-level patterns that reveal a mechanistic signature of difficulty: easy-to-unlearn samples are associated with shorter, shallower interactions concentrated in earlier-to-intermediate parts of the original model, whereas hard samples rely on longer and deeper pathways closer to late-stage computation. Compared to existing qualitative studies, CUD takes a first step toward a principled, fine-grained, and interpretable analysis of unlearning difficulty; and motivates the development of unlearning methods grounded in model mechanisms.
{
"annotation_id": "20d8353c-6ebc-48bd-8555-889113530a96",
"date_created": "2026-02-17T05:53:19.959000Z",
"date_modified": "2026-02-17T05:53:19.959000Z",
"file_hash": "0cc253ab0ce5ebd7d70913b3acc764279b1c44ee3388020ab63f40652012f023",
"private": false,
"record": {
"abstract": "Machine unlearning is becoming essential for building trustworthy and compliant language models. Yet unlearning success varies considerably across individual samples: some are reliably erased, while others persist despite the same procedure. We argue that this disparity is not only a data-side phenomenon, but also reflects model-internal mechanisms that encode and protect memorized information. We study this problem from a mechanistic perspective based on model circuits--structured interaction pathways that govern how predictions are formed. We propose Circuit-guided Unlearning Difficulty (CUD), a {\\em pre-unlearning} metric that assigns each sample a continuous difficulty score using circuit-level signals. Extensive experiments demonstrate that CUD reliably separates intrinsically easy and hard samples, and remains stable across unlearning methods. We identify key circuit-level patterns that reveal a mechanistic signature of difficulty: easy-to-unlearn samples are associated with shorter, shallower interactions concentrated in earlier-to-intermediate parts of the original model, whereas hard samples rely on longer and deeper pathways closer to late-stage computation. Compared to existing qualitative studies, CUD takes a first step toward a principled, fine-grained, and interpretable analysis of unlearning difficulty; and motivates the development of unlearning methods grounded in model mechanisms.",
"arxiv_id": "2601.09624",
"authors": [
"Jiali Cheng",
"Ziheng Chen",
"Chirag Agarwal",
"Hadi Amiri"
],
"categories": [
"cs.LG",
"cs.AI",
"cs.CL",
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Toward Understanding Unlearning Difficulty: A Mechanistic Perspective and Circuit-Guided Difficulty Metric",
"url": "https://arxiv.org/abs/2601.09624",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "8cb41fc6-f7eb-422b-8529-d1355e922c65",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}