Score: 0

Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL

Published: January 14, 2026 | arXiv ID: 2601.09876v1

By: Yifei Shen , Yilun Zhao , Justice Ou and more

Real-world clinical text-to-SQL requires reasoning over heterogeneous EHR tables, temporal windows, and patient-similarity cohorts to produce executable queries. We introduce CLINSQL, a benchmark of 633 expert-annotated tasks on MIMIC-IV v3.1 that demands multi-table joins, clinically meaningful filters, and executable SQL. Solving CLINSQL entails navigating schema metadata and clinical coding systems, handling long contexts, and composing multi-step queries beyond traditional text-to-SQL. We evaluate 22 proprietary and open-source models under Chain-of-Thought self-refinement and use rubric-based SQL analysis with execution checks that prioritize critical clinical requirements. Despite recent advances, performance remains far from clinical reliability: on the test set, GPT-5-mini attains 74.7% execution score, DeepSeek-R1 leads open-source at 69.2% and Gemini-2.5-Pro drops from 85.5% on Easy to 67.2% on Hard. Progress on CLINSQL marks tangible advances toward clinically reliable text-to-SQL for real-world EHR analytics.

Generating patient cohorts from electronic health records using two-step retrieval-augmented text-to-SQL generation

Computation and Language

Finds sick people for medical studies faster.

28 Feb 2025 1

90%

BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases

Computation and Language

Helps computers answer science questions from data.

23 May 2025 1

90%

Generating patient cohorts from electronic health records using two-step retrieval-augmented text-to-SQL generation

Computation and Language

Finds the right patients for medical studies faster.

28 Feb 2025 0

View PDF Login to Bookmark

Patient-Similarity Cohort Reasoning in Clinical Text-to-SQL

Technical Abstract

Generating patient cohorts from electronic health records using two-step retrieval-augmented text-to-SQL generation

BiomedSQL: Text-to-SQL for Scientific Reasoning on Biomedical Knowledge Bases

Generating patient cohorts from electronic health records using two-step retrieval-augmented text-to-SQL generation