dorsal/arxiv
View SchemaKnowledge-based learning in Text-RAG and Image-RAG
| Authors | Alexander Shim, Khalil Saieh, Samuel Clarke |
|---|---|
| Categories | |
| ArXiv ID | 2601.08226vv1 |
| URL | https://arxiv.org/abs/2601.08226 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
This research analyzed and compared the multi-modal approach in the Vision Transformer(EVA-ViT) based image encoder with the LlaMA or ChatGPT LLM to reduce the hallucination problem and detect diseases in chest x-ray images. In this research, we utilized the NIH Chest X-ray image to train the model and compared it in image-based RAG, text-based RAG, and baseline. [3] [5] In a result, the text-based RAG[2] e!ectively reduces the hallucination problem by using external knowledge information, and the image-based RAG improved the prediction con"dence and calibration by using the KNN methods. [4] Moreover, the GPT LLM showed better performance, a low hallucination rate, and better Expected Calibration Error(ECE) than Llama Llama-based model. This research shows the challenge of data imbalance, a complex multi-stage structure, but suggests a large experience environment and a balanced example of use.
{
"annotation_id": "9d312548-457c-4d6b-8def-f262f07248b1",
"date_created": "2026-02-17T05:53:16.174000Z",
"date_modified": "2026-02-17T05:53:16.174000Z",
"file_hash": "3f9c01b95ec858e54941cec15b265b6ff09dfc2d829bdc9e37bcab74b8303a91",
"private": false,
"record": {
"abstract": "This research analyzed and compared the multi-modal approach in the Vision Transformer(EVA-ViT) based image encoder with the LlaMA or ChatGPT LLM to reduce the hallucination problem and detect diseases in chest x-ray images. In this research, we utilized the NIH Chest X-ray image to train the model and compared it in image-based RAG, text-based RAG, and baseline. [3] [5] In a result, the text-based RAG[2] e!ectively reduces the hallucination problem by using external knowledge information, and the image-based RAG improved the prediction con\"dence and calibration by using the KNN methods. [4] Moreover, the GPT LLM showed better performance, a low hallucination rate, and better Expected Calibration Error(ECE) than Llama Llama-based model. This research shows the challenge of data imbalance, a complex multi-stage structure, but suggests a large experience environment and a balanced example of use.",
"arxiv_id": "2601.08226",
"authors": [
"Alexander Shim",
"Khalil Saieh",
"Samuel Clarke"
],
"categories": [
"cs.CV",
"cs.AI"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Knowledge-based learning in Text-RAG and Image-RAG",
"url": "https://arxiv.org/abs/2601.08226",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "b389ac6e-1fc5-4226-9517-e8ee6dfa6b70",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}