dorsal/arxiv
View SchemaMedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation
| Authors | Yuelyu Ji, Min Gu Kwak, Hang Zhang, Xizhi Wu, Chenyu Li, Yanshan Wang |
|---|---|
| Categories | |
| ArXiv ID | 2601.06519vv1 |
| URL | https://arxiv.org/abs/2601.06519 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications. We introduce MedRAGChecker, a claim-level verification and diagnostic framework for biomedical RAG. Given a question, retrieved evidence, and a generated answer, MedRAGChecker decomposes the answer into atomic claims and estimates claim support by combining evidence-grounded natural language inference (NLI) with biomedical knowledge-graph (KG) consistency signals. Aggregating claim decisions yields answer-level diagnostics that help disentangle retrieval and generation failures, including faithfulness, under-evidence, contradiction, and safety-critical error rates. To enable scalable evaluation, we distill the pipeline into compact biomedical models and use an ensemble verifier with class-specific reliability weighting. Experiments on four biomedical QA benchmarks show that MedRAGChecker reliably flags unsupported and contradicted claims and reveals distinct risk profiles across generators, particularly on safety-critical biomedical relations.
{
"annotation_id": "59540b5b-340c-4712-a2c5-3a192b0db839",
"date_created": "2026-02-17T05:53:08.709000Z",
"date_modified": "2026-02-17T05:53:08.709000Z",
"file_hash": "aa1d1f0a705359f6abccabad0b4dea1ed7ca148d56ad5ab4d3f1054f3e63e53f",
"private": false,
"record": {
"abstract": "Biomedical retrieval-augmented generation (RAG) can ground LLM answers in medical literature, yet long-form outputs often contain isolated unsupported or contradictory claims with safety implications.\n We introduce MedRAGChecker, a claim-level verification and diagnostic framework for biomedical RAG.\n Given a question, retrieved evidence, and a generated answer, MedRAGChecker decomposes the answer into atomic claims and estimates claim support by combining evidence-grounded natural language inference (NLI) with biomedical knowledge-graph (KG) consistency signals.\n Aggregating claim decisions yields answer-level diagnostics that help disentangle retrieval and generation failures, including faithfulness, under-evidence, contradiction, and safety-critical error rates.\n To enable scalable evaluation, we distill the pipeline into compact biomedical models and use an ensemble verifier with class-specific reliability weighting.\n Experiments on four biomedical QA benchmarks show that MedRAGChecker reliably flags unsupported and contradicted claims and reveals distinct risk profiles across generators, particularly on safety-critical biomedical relations.",
"arxiv_id": "2601.06519",
"authors": [
"Yuelyu Ji",
"Min Gu Kwak",
"Hang Zhang",
"Xizhi Wu",
"Chenyu Li",
"Yanshan Wang"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "MedRAGChecker: Claim-Level Verification for Biomedical Retrieval-Augmented Generation",
"url": "https://arxiv.org/abs/2601.06519",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "573b92f3-84eb-42dc-b3d7-b82a7cc09383",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}