GraPPa: Grammar-Augmented Pre-Training for Table Semantic Parsing
arXiv:2009.13845
Abstract
We present GraPPa, an effective pre-training approach for table semantic parsing that learns a compositional inductive bias in the joint representations of textual and tabular data. We construct synthetic question-SQL pairs over high-quality tables via a synchronous context-free grammar (SCFG) induced from existing text-to-SQL datasets. We pre-train our model on the synthetic data using a novel text-schema linking objective that predicts the syntactic role of a table field in the SQL for each question-SQL pair. To maintain the model's ability to represent real-world data, we also include masked language modeling (MLM) over several existing table-and-language datasets to regularize the pre-training process. On four popular fully supervised and weakly supervised table semantic parsing benchmarks, GraPPa significantly outperforms RoBERTa-large as the feature representation layers and establishes new state-of-the-art results on all of them.
16 pages; Accepted to ICLR 2021
References in corpus (20)
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Learning to Map Sentences to Logical Form: Structured Classification with Probabilistic Categorial Grammars
- Seq2SQL: Generating Structured Queries from Natural Language using Reinforcement Learning
- TAPAS: Weakly Supervised Table Parsing via Pre-training
- Layer Normalization
- SQLNet: Generating Structured Queries From Natural Language Without Reinforcement Learning
- TabFact: A Large-scale Dataset for Table-based Fact Verification
- A Comprehensive Exploration on WikiSQL with Table-Aware Word Contextualization
- Table2Vec: Neural Word and Entity Embeddings for Table Population and Retrieval
- Pre-training via Paraphrasing
- Learning to Generalize from Sparse and Underspecified Rewards
- IncSQL: Training Incremental Text-to-SQL Parsers with Non-Deterministic Oracles
- TURL: Table Understanding through Representation Learning
- X-SQL: reinforce schema representation with context
- Hybrid Ranking Network for Text-to-SQL
- Content Enhanced BERT-based Text-to-SQL Generation
- Towards Complex Text-to-SQL in Cross-Domain Database with Intermediate Representation
- A Discrete Hard EM Approach for Weakly Supervised Question Answering
- HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data
- Logical Natural Language Generation from Open-Domain Tables
Cited by in corpus (16)
- TAPEX: Table Pre-training via Learning a Neural SQL Executor
- GP: Context-free Grammar Pre-training for Text-to-SQL Parsers
- Relation Aware Semi-autoregressive Semantic Parsing for NL2SQL
- FeTaQA: Free-form Table Question Answering
- Optimizing Deeper Transformers on Small Datasets
- Decoupled Dialogue Modeling and Semantic Parsing for Multi-Turn Text-to-SQL
- Data Augmentation with Hierarchical SQL-to-Question Generation for Cross-domain Text-to-SQL Parsing
- Constrained Language Models Yield Few-Shot Semantic Parsers
- MT-Teql: Evaluating and Augmenting Consistency of Text-to-SQL Models with Metamorphic Testing
- End-to-End Cross-Domain Text-to-SQL Semantic Parsing with Auxiliary Task
- ShadowGNN: Graph Projection Neural Network for Text-to-SQL Parser
- mRAT-SQL+GAP:A Portuguese Text-to-SQL Transformer
- Learning to Synthesize Data for Semantic Parsing
- KaggleDBQA: Realistic Evaluation of Text-to-SQL Parsers
- SPARQLing Database Queries from Intermediate Question Decompositions
- Logic-level Evidence Retrieval and Graph-based Verification Network for Table-based Fact Verification