dorsal/arxiv
View SchemaT3: Benchmarking Sycophancy and Skepticism in Causal Judgment
| Authors | Edward Y. Chang |
|---|---|
| Categories | |
| ArXiv ID | 2601.08258vv2 |
| URL | https://arxiv.org/abs/2601.08258 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
We introduce T3 (Testing Trustworthy Thinking), a diagnostic benchmark designed to rigorously evaluate LLM causal judgment across Pearl's Ladder of Causality. Comprising 454 expert-curated vignettes, T3 prioritizes high-resolution failure analysis, decomposing performance into Utility (sensitivity), Safety (specificity), and Wise Refusal on underdetermined cases. By applying T3 to frontier models, we diagnose two distinct pathologies: a "Skepticism Trap" at L1 (where safety-tuned models like Claude Haiku reject 60% of valid links) and a non-monotonic Scaling Paradox at L3. In the latter, the larger GPT-5.2 underperforms GPT-4-Turbo by 55 points on ambiguous counterfactuals, driven by a collapse into paralysis (excessive hedging) rather than hallucination. Finally, we use the benchmark to validate a process-verified protocol (RCA), showing that T3 successfully captures the restoration of decisive causal judgment under structured verification.
{
"annotation_id": "2c720be4-54cf-468b-a8e7-86817e8ccb29",
"date_created": "2026-02-17T05:53:15.955000Z",
"date_modified": "2026-02-17T05:53:15.955000Z",
"file_hash": "96270c05eed2944ed80007af53ce47baed5b8ba08ebd2704b3eb344a7a48db39",
"private": false,
"record": {
"abstract": "We introduce T3 (Testing Trustworthy Thinking), a diagnostic benchmark designed to rigorously evaluate LLM causal judgment across Pearl\u0027s Ladder of Causality. Comprising 454 expert-curated vignettes, T3 prioritizes high-resolution failure analysis, decomposing performance into Utility (sensitivity), Safety (specificity), and Wise Refusal on underdetermined cases. By applying T3 to frontier models, we diagnose two distinct pathologies: a \"Skepticism Trap\" at L1 (where safety-tuned models like Claude Haiku reject 60% of valid links) and a non-monotonic Scaling Paradox at L3. In the latter, the larger GPT-5.2 underperforms GPT-4-Turbo by 55 points on ambiguous counterfactuals, driven by a collapse into paralysis (excessive hedging) rather than hallucination. Finally, we use the benchmark to validate a process-verified protocol (RCA), showing that T3 successfully captures the restoration of decisive causal judgment under structured verification.",
"arxiv_id": "2601.08258",
"authors": [
"Edward Y. Chang"
],
"categories": [
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "T3: Benchmarking Sycophancy and Skepticism in Causal Judgment",
"url": "https://arxiv.org/abs/2601.08258",
"version": "v2"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "53f706b3-2c2d-4b09-97f3-fcd59cc225c7",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}