A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?
arXiv:2408.05109 · doi:10.1109/TKDE.2025.3592032
Abstract
Translating users' natural language queries (NL) into SQL queries (i.e., Text-to-SQL, a.k.a. NL2SQL) can significantly reduce barriers to accessing relational databases and support various commercial applications. The performance of Text-to-SQL has been greatly enhanced with the emergence of Large Language Models (LLMs). In this survey, we provide a comprehensive review of Text-to-SQL techniques powered by LLMs, covering its entire lifecycle from the following four aspects: (1) Model: Text-to-SQL translation techniques that tackle not only NL ambiguity and under-specification, but also properly map NL with database schema and instances; (2) Data: From the collection of training data, data synthesis due to training data scarcity, to Text-to-SQL benchmarks; (3) Evaluation: Evaluating Text-to-SQL methods from multiple angles using different metrics and granularities; and (4) Error Analysis: analyzing Text-to-SQL errors to find the root cause and guiding Text-to-SQL models to evolve. Moreover, we offer a rule of thumb for developing Text-to-SQL solutions. Finally, we discuss the research challenges and open problems of Text-to-SQL in the LLMs era. Text-to-SQL Handbook: https://github.com/HKUSTDial/NL2SQL_Handbook
20 pages, 11 figures, 3 tables
References in corpus (6)
- The Effect of Sampling Temperature on Problem Solving in Large Language Models
- Towards Natural Language Interfaces for Data Visualization: A Survey
- Structure-Grounded Pretraining for Text-to-SQL
- The Dawn of Natural Language to SQL: Are We Fully Ready?
- Metasql: A Generate-then-Rank Framework for Natural Language to SQL Translation
- NL2SQL-BUGs: A Benchmark for Detecting Semantic Errors in NL2SQL Translation