Hybrid Ranking Network for Text-to-SQL
arXiv:2008.04759
Abstract
In this paper, we study how to leverage pre-trained language models in Text-to-SQL. We argue that previous approaches under utilize the base language models by concatenating all columns together with the NL question and feeding them into the base language model in the encoding stage. We propose a neat approach called Hybrid Ranking Network (HydraNet) which breaks down the problem into column-wise ranking and decoding and finally assembles the column-wise outputs into a SQL query by straightforward rules. In this approach, the encoder is given a NL question and one individual column, which perfectly aligns with the original tasks BERT/RoBERTa is trained on, and hence we avoid any ad-hoc pooling or additional encoding layers which are necessary in prior approaches. Experiments on the WikiSQL dataset show that the proposed approach is very effective, achieving the top place on the leaderboard.
References in corpus (3)
Cited by in corpus (7)
- Translating Place-Related Questions to GeoSPARQL Queries
- Application of Deep Learning in Generating Structured Radiology Reports: A Transformer-Based Technique
- Bridging Textual and Tabular Data for Cross-Domain Text-to-SQL Semantic Parsing
- Semantic Evaluation for Text-to-SQL with Distilled Test Suites
- SeqGenSQL -- A Robust Sequence Generation Model for Structured Query Language
- Query Understanding for Natural Language Enterprise Search
- A Mutual Information Maximization Approach for the Spurious Solution Problem in Weakly Supervised Question Answering