dorsal/arxiv
View SchemaInterpretable Text Classification Applied to the Detection of LLM-generated Creative Writing
| Authors | Minerva Suvanto, Andrea McGlinchey, Mattias Wahde, Peter J Barclay |
|---|---|
| Categories | |
| ArXiv ID | 2601.07368vv1 |
| URL | https://arxiv.org/abs/2601.07368 |
| License | http://creativecommons.org/licenses/by-nc-nd/4.0/ |
Abstract
We consider the problem of distinguishing human-written creative fiction (excerpts from novels) from similar text generated by an LLM. Our results show that, while human observers perform poorly (near chance levels) on this binary classification task, a variety of machine-learning models achieve accuracy in the range 0.93 - 0.98 over a previously unseen test set, even using only short samples and single-token (unigram) features. We therefore employ an inherently interpretable (linear) classifier (with a test accuracy of 0.98), in order to elucidate the underlying reasons for this high accuracy. In our analysis, we identify specific unigram features indicative of LLM-generated text, one of the most important being that the LLM tends to use a larger variety of synonyms, thereby skewing the probability distributions in a manner that is easy to detect for a machine learning classifier, yet very difficult for a human observer. Four additional explanation categories were also identified, namely, temporal drift, Americanisms, foreign language usage, and colloquialisms. As identification of the AI-generated text depends on a constellation of such features, the classification appears robust, and therefore not easy to circumvent by malicious actors intent on misrepresenting AI-generated text as human work.
{
"annotation_id": "7ea0c63a-6075-4ac6-8789-972077789429",
"date_created": "2026-02-17T05:53:11.819000Z",
"date_modified": "2026-02-17T05:53:11.819000Z",
"file_hash": "6ea181a5b47738d31a97f4143be38cad50b1c523d3c7c9589540a7670cb5fc4c",
"private": false,
"record": {
"abstract": "We consider the problem of distinguishing human-written creative fiction (excerpts from novels) from similar text generated by an LLM. Our results show that, while human observers perform poorly (near chance levels) on this binary classification task, a variety of machine-learning models achieve accuracy in the range 0.93 - 0.98 over a previously unseen test set, even using only short samples and single-token (unigram) features. We therefore employ an inherently interpretable (linear) classifier (with a test accuracy of 0.98), in order to elucidate the underlying reasons for this high accuracy. In our analysis, we identify specific unigram features indicative of LLM-generated text, one of the most important being that the LLM tends to use a larger variety of synonyms, thereby skewing the probability distributions in a manner that is easy to detect for a machine learning classifier, yet very difficult for a human observer. Four additional explanation categories were also identified, namely, temporal drift, Americanisms, foreign language usage, and colloquialisms. As identification of the AI-generated text depends on a constellation of such features, the classification appears robust, and therefore not easy to circumvent by malicious actors intent on misrepresenting AI-generated text as human work.",
"arxiv_id": "2601.07368",
"authors": [
"Minerva Suvanto",
"Andrea McGlinchey",
"Mattias Wahde",
"Peter J Barclay"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by-nc-nd/4.0/",
"title": "Interpretable Text Classification Applied to the Detection of LLM-generated Creative Writing",
"url": "https://arxiv.org/abs/2601.07368",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "516158a3-5dcf-4e73-bb15-7975b51388ee",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}