dorsal/arxiv
View SchemaContext-Aware Decoding for Faithful Vision-Language Generation
| Authors | Mehrdad Fazli, Bowen Wei, Ziwei Zhu |
|---|---|
| Categories | |
| ArXiv ID | 2601.05939vv1 |
| URL | https://arxiv.org/abs/2601.05939 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Hallucinations, generating responses inconsistent with the visual input, remain a critical limitation of large vision-language models (LVLMs), especially in open-ended tasks such as image captioning and visual reasoning. In this work, we probe the layer-wise generation dynamics that drive hallucinations and propose a training-free mitigation strategy. Employing the Logit Lens, we examine how LVLMs construct next-token distributions across decoder layers, uncovering a pronounced commitment-depth gap: truthful tokens accumulate probability mass on their final candidates earlier than hallucinatory ones. Drawing on this discovery, we introduce Context Embedding Injection (CEI), a lightweight method that harnesses the hidden state of the last input token-the context embedding-as a grounding signal to maintain visual fidelity throughout decoding and curb hallucinations. Evaluated on the CHAIR, AMBER, and MMHal-Bench benchmarks (with a maximum token length of 512), CEI outperforms state-of-the-art baselines across three LVLMs, with its dynamic variant yielding the lowest overall hallucination rates. By integrating novel mechanistic insights with a scalable intervention, this work advances the mitigation of hallucinations in LVLMs.
{
"annotation_id": "d48f821e-4cd6-4915-92c3-ff9acc032d43",
"date_created": "2026-02-17T05:53:04.951000Z",
"date_modified": "2026-02-17T05:53:04.951000Z",
"file_hash": "4e21509daffb8138c24b13bf84a1547f090c32c973a44eec67d3ecc927df5be3",
"private": false,
"record": {
"abstract": "Hallucinations, generating responses inconsistent with the visual input, remain a critical limitation of large vision-language models (LVLMs), especially in open-ended tasks such as image captioning and visual reasoning. In this work, we probe the layer-wise generation dynamics that drive hallucinations and propose a training-free mitigation strategy. Employing the Logit Lens, we examine how LVLMs construct next-token distributions across decoder layers, uncovering a pronounced commitment-depth gap: truthful tokens accumulate probability mass on their final candidates earlier than hallucinatory ones. Drawing on this discovery, we introduce Context Embedding Injection (CEI), a lightweight method that harnesses the hidden state of the last input token-the context embedding-as a grounding signal to maintain visual fidelity throughout decoding and curb hallucinations. Evaluated on the CHAIR, AMBER, and MMHal-Bench benchmarks (with a maximum token length of 512), CEI outperforms state-of-the-art baselines across three LVLMs, with its dynamic variant yielding the lowest overall hallucination rates. By integrating novel mechanistic insights with a scalable intervention, this work advances the mitigation of hallucinations in LVLMs.",
"arxiv_id": "2601.05939",
"authors": [
"Mehrdad Fazli",
"Bowen Wei",
"Ziwei Zhu"
],
"categories": [
"cs.CV"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "Context-Aware Decoding for Faithful Vision-Language Generation",
"url": "https://arxiv.org/abs/2601.05939",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6e8a861d-44ee-432d-a9ab-0d96efdb2e10",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}