dorsal/arxiv
View SchemaHot-Start from Pixels: Low-Resolution Visual Tokens for Chinese Language Modeling
| Authors | Shuyang Xiang, Hao Guan |
|---|---|
| Categories | |
| ArXiv ID | 2601.09566vv1 |
| URL | https://arxiv.org/abs/2601.09566 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Large language models typically represent Chinese characters as discrete index-based tokens, largely ignoring their visual form. For logographic scripts, visual structure carries semantic and phonetic information, which may aid prediction. We investigate whether low-resolution visual inputs can serve as an alternative for character-level modeling. Instead of token IDs, our decoder receives grayscale images of individual characters, with resolutions as low as $8 \times 8$ pixels. Remarkably, these inputs achieve 39.2\% accuracy, comparable to the index-based baseline of 39.1\%. Such low-resource settings also exhibit a pronounced \emph{hot-start} effect: by 0.4\% of total training, accuracy reaches above 12\%, while index-based models lag at below 6\%. Overall, our results demonstrate that minimal visual structure can provide a robust and efficient signal for Chinese language modeling, offering an alternative perspective on character representation that complements traditional index-based approaches.
{
"annotation_id": "67c52983-8cb7-45bb-bbcc-34d18b6cecf5",
"date_created": "2026-02-17T05:53:19.964000Z",
"date_modified": "2026-02-17T05:53:19.964000Z",
"file_hash": "b3b7172aeb321f902d64ad5a8a34fa3235581a8a46bd3b0f851fb0a481c40fad",
"private": false,
"record": {
"abstract": "Large language models typically represent Chinese characters as discrete index-based tokens, largely ignoring their visual form. For logographic scripts, visual structure carries semantic and phonetic information, which may aid prediction. We investigate whether low-resolution visual inputs can serve as an alternative for character-level modeling. Instead of token IDs, our decoder receives grayscale images of individual characters, with resolutions as low as $8 \\times 8$ pixels. Remarkably, these inputs achieve 39.2\\% accuracy, comparable to the index-based baseline of 39.1\\%. Such low-resource settings also exhibit a pronounced \\emph{hot-start} effect: by 0.4\\% of total training, accuracy reaches above 12\\%, while index-based models lag at below 6\\%. Overall, our results demonstrate that minimal visual structure can provide a robust and efficient signal for Chinese language modeling, offering an alternative perspective on character representation that complements traditional index-based approaches.",
"arxiv_id": "2601.09566",
"authors": [
"Shuyang Xiang",
"Hao Guan"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Hot-Start from Pixels: Low-Resolution Visual Tokens for Chinese Language Modeling",
"url": "https://arxiv.org/abs/2601.09566",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "88284f72-f286-4566-be9f-1e052c2d72fd",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}