dorsal/arxiv
View SchemaPatient-Similarity Cohort Reasoning in Clinical Text-to-SQL
| Authors | Yifei Shen, Yilun Zhao, Justice Ou, Tinglin Huang, Arman Cohan |
|---|---|
| Categories | |
| ArXiv ID | 2601.09876vv1 |
| URL | https://arxiv.org/abs/2601.09876 |
| License | http://arxiv.org/licenses/nonexclusive-distrib/1.0/ |
Abstract
Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce CLINSQL, a benchmark of 633 expert-annotated tasks on MIMIC-IV v3.1 that demands multi-table joins, clinically meaningful filters, and executable SQL. Solving CLINSQL entails navigating schema metadata and clinical coding systems, handling long contexts, and composing multi-step queries beyond traditional text-to-SQL. We evaluate 22 proprietary and open-source models under Chain-of-Thought self-refinement and use rubric-based SQL analysis with execution checks that prioritize critical clinical requirements. Despite recent advances, performance remains far from clinical reliability: on the test set, GPT-5-mini attains 74.7% execution score, DeepSeek-R1 leads open-source at 69.2% and Gemini-2.5-Pro drops from 85.5% on Easy to 67.2% on Hard. Progress on CLINSQL marks tangible advances toward clinically reliable text-to-SQL for real-world EHR analytics.
{
"annotation_id": "28383bfb-b462-4072-b9cf-033bfac63996",
"date_created": "2026-02-17T05:53:23.750000Z",
"date_modified": "2026-02-17T05:53:23.750000Z",
"file_hash": "3be03faedb0446eead1566aaf044bc2470a092c96d381d069c5dbf1e9b24fbac",
"private": false,
"record": {
"abstract": "Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce CLINSQL, a benchmark of 633 expert-annotated tasks on MIMIC-IV v3.1 that demands multi-table joins, clinically meaningful filters, and executable SQL. Solving CLINSQL entails navigating schema metadata and clinical coding systems, handling long contexts, and composing multi-step queries beyond traditional text-to-SQL. We evaluate 22 proprietary and open-source models under Chain-of-Thought self-refinement and use rubric-based SQL analysis with execution checks that prioritize critical clinical requirements. Despite recent advances, performance remains far from clinical reliability: on the test set, GPT-5-mini attains 74.7% execution score, DeepSeek-R1 leads open-source at 69.2% and Gemini-2.5-Pro drops from 85.5% on Easy to 67.2% on Hard. Progress on CLINSQL marks tangible advances toward clinically reliable text-to-SQL for real-world EHR analytics.",
"arxiv_id": "2601.09876",
"authors": [
"Yifei Shen",
"Yilun Zhao",
"Justice Ou",
"Tinglin Huang",
"Arman Cohan"
],
"categories": [
"cs.CL"
],
"license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
"title": "Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL",
"url": "https://arxiv.org/abs/2601.09876",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "b74769ab-a809-482c-a05a-c6420f6a5972",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}