dorsal/arxiv
View SchemaEntropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM
| Authors | Pedro Memoli Buffa, Luciano Del Corro |
|---|---|
| Categories | |
| ArXiv ID | 2601.09001vv1 |
| URL | https://arxiv.org/abs/2601.09001 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Deploying LLMs raises two coupled challenges: (1) monitoring - estimating where a model underperforms as traffic and domains drift - and (2) improvement - prioritizing data acquisition to close the largest performance gaps. We test whether an inference-time signal can estimate slice-level accuracy under domain shift. For each response, we compute an output-entropy profile from final-layer next-token probabilities (from top-k logprobs) and summarize it with eleven statistics. A lightweight classifier predicts instance correctness, and averaging predicted probabilities yields a domain-level accuracy estimate. We evaluate on ten STEM reasoning benchmarks with exhaustive train/test compositions (k in {1,2,3,4}; all "10 choose k" combinations), across nine LLMs from six families (3B-20B). Estimates often track held-out benchmark accuracy, and several models show near-monotonic ordering of domains. Output-entropy profiles are thus an accessible signal for scalable monitoring and for targeting data acquisition.
{
"annotation_id": "bae15f26-d88a-48aa-b155-9a0530a4dce0",
"date_created": "2026-02-17T05:53:20.186000Z",
"date_modified": "2026-02-17T05:53:20.186000Z",
"file_hash": "290d8a3beb08080ff4be95731116e3835fc02b9759f9d3fafa28a7b183da8d3f",
"private": false,
"record": {
"abstract": "Deploying LLMs raises two coupled challenges: (1) monitoring - estimating where a model underperforms as traffic and domains drift - and (2) improvement - prioritizing data acquisition to close the largest performance gaps. We test whether an inference-time signal can estimate slice-level accuracy under domain shift. For each response, we compute an output-entropy profile from final-layer next-token probabilities (from top-k logprobs) and summarize it with eleven statistics. A lightweight classifier predicts instance correctness, and averaging predicted probabilities yields a domain-level accuracy estimate. We evaluate on ten STEM reasoning benchmarks with exhaustive train/test compositions (k in {1,2,3,4}; all \"10 choose k\" combinations), across nine LLMs from six families (3B-20B). Estimates often track held-out benchmark accuracy, and several models show near-monotonic ordering of domains. Output-entropy profiles are thus an accessible signal for scalable monitoring and for targeting data acquisition.",
"arxiv_id": "2601.09001",
"authors": [
"Pedro Memoli Buffa",
"Luciano Del Corro"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Entropy Sentinel: Continuous LLM Accuracy Monitoring from Decoding Entropy Traces in STEM",
"url": "https://arxiv.org/abs/2601.09001",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "93fc5eed-431e-4c10-b117-90912dc884e5",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}