dorsal/arxiv
View SchemaCSR-RAG: An Efficient Retrieval System for Text-to-SQL on the Enterprise Scale
| Authors | Rajpreet Singh, Novak Boškov, Lawrence Drabeck, Aditya Gudal, Manzoor A. Khan |
|---|---|
| Categories | |
| ArXiv ID | 2601.06564vv1 |
| URL | https://arxiv.org/abs/2601.06564 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Natural language to SQL translation (Text-to-SQL) is one of the long-standing problems that has recently benefited from advances in Large Language Models (LLMs). While most academic Text-to-SQL benchmarks request schema description as a part of natural language input, enterprise-scale applications often require table retrieval before SQL query generation. To address this need, we propose a novel hybrid Retrieval Augmented Generation (RAG) system consisting of contextual, structural, and relational retrieval (CSR-RAG) to achieve computationally efficient yet sufficiently accurate retrieval for enterprise-scale databases. Through extensive enterprise benchmarks, we demonstrate that CSR-RAG achieves up to 40% precision and over 80% recall while incurring a negligible average query generation latency of only 30ms on commodity data center hardware, which makes it appropriate for modern LLM-based enterprise-scale systems.
{
"annotation_id": "fc4f72c8-d082-4099-9c23-e09948a8881e",
"date_created": "2026-02-17T05:53:07.765000Z",
"date_modified": "2026-02-17T05:53:07.765000Z",
"file_hash": "54432c24e75ea0b23890ab3b52b4b7c8ae769e372d8d34da73fa6e8206cfc91c",
"private": false,
"record": {
"abstract": "Natural language to SQL translation (Text-to-SQL) is one of the long-standing problems that has recently benefited from advances in Large Language Models (LLMs). While most academic Text-to-SQL benchmarks request schema description as a part of natural language input, enterprise-scale applications often require table retrieval before SQL query generation. To address this need, we propose a novel hybrid Retrieval Augmented Generation (RAG) system consisting of contextual, structural, and relational retrieval (CSR-RAG) to achieve computationally efficient yet sufficiently accurate retrieval for enterprise-scale databases. Through extensive enterprise benchmarks, we demonstrate that CSR-RAG achieves up to 40% precision and over 80% recall while incurring a negligible average query generation latency of only 30ms on commodity data center hardware, which makes it appropriate for modern LLM-based enterprise-scale systems.",
"arxiv_id": "2601.06564",
"authors": [
"Rajpreet Singh",
"Novak Bo\u0161kov",
"Lawrence Drabeck",
"Aditya Gudal",
"Manzoor A. Khan"
],
"categories": [
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "CSR-RAG: An Efficient Retrieval System for Text-to-SQL on the Enterprise Scale",
"url": "https://arxiv.org/abs/2601.06564",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "dc490b72-8a71-41b1-8f91-c182565ac768",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}