Structure-Grounded Pretraining for Text-to-SQL
arXiv:2010.12773 · doi:10.18653/v1/2021.naacl-main.105
Abstract
Learning to capture text-table alignment is essential for tasks like text-to-SQL. A model needs to correctly recognize natural language references to columns and values and to ground them in the given database schema. In this paper, we present a novel weakly supervised Structure-Grounded pretraining framework (StruG) for text-to-SQL that can effectively learn to capture text-table alignment based on a parallel text-table corpus. We identify a set of novel prediction tasks: column grounding, value grounding and column-value mapping, and leverage them to pretrain a text-table encoder. Additionally, to evaluate different methods under more realistic text-table alignment settings, we create a new evaluation set Spider-Realistic based on Spider dev set with explicit mentions of column names removed, and adopt eight existing text-to-SQL datasets for cross-database evaluation. STRUG brings significant improvement over BERT-LARGE in all settings. Compared with existing pretraining methods such as GRAPPA, STRUG achieves similar performance on Spider, and outperforms all baselines on more realistic sets. The Spider-Realistic dataset is available at https://doi.org/10.5281/zenodo.5205322.
Accepted to NAACL 2021. The Spider-Realistic dataset is available at https://doi.org/10.5281/zenodo.5205322
Cited by in corpus (6)
- The Dawn of Natural Language to SQL: Are We Fully Ready?
- A Survey of Text-to-SQL in the Era of LLMs: Where are we, and where are we going?
- Grounding Natural Language to SQL Translation with Data-Based Self-Explanations
- DataVisT5: A Pre-trained Language Model for Jointly Understanding Text and Data Visualization
- Multi-Turn Interactions for Text-to-SQL with Large Language Models
- LEAP: LLM-powered End-to-end Automatic Library for Processing Social Science Queries on Unstructured Data