Unlocking Compositional Generalization in Pre-trained Models Using Intermediate Representations
arXiv:2104.07478
Abstract
Sequence-to-sequence (seq2seq) models are prevalent in semantic parsing, but have been found to struggle at out-of-distribution compositional generalization. While specialized model architectures and pre-training of seq2seq models have been proposed to address this issue, the former often comes at the cost of generality and the latter only shows limited success. In this paper, we study the impact of intermediate representations on compositional generalization in pre-trained seq2seq models, without changing the model architecture at all, and identify key aspects for designing effective representations. Instead of training to directly map natural language to an executable form, we map to a reversible or lossy intermediate representation that has stronger structural correspondence with natural language. The combination of our proposed intermediate representations and pre-trained models is surprisingly effective, where the best combinations obtain a new state-of-the-art on CFQ (+14.8 accuracy points) and on the template-splits of three text-to-SQL datasets (+15.0 to +19.4 accuracy points). This work highlights that intermediate representations provide an important and potentially overlooked degree of freedom for improving the compositional generalization abilities of pre-trained seq2seq models.
References in corpus (8)
- Learning to Map Sentences to Logical Form: Structured Classification with Probabilistic Categorial Grammars
- Improving Text-to-SQL Evaluation Methodology
- The Evolved Transformer
- Compositional generalization in a deep seq2seq model by separating syntax and semantics
- Compositional Generalization in Semantic Parsing: Pre-training vs. Specialized Architectures
- Learning to Infer Program Sketches
- Learning Compositional Rules via Neural Program Synthesis
- Compositional Generalization via Semantic Tagging
Cited by in corpus (9)
- Modern Baselines for SPARQL Semantic Parsing
- ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
- The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers
- Constrained Language Models Yield Few-Shot Semantic Parsers
- Inducing Transformer's Compositional Generalization Ability via Auxiliary Sequence Prediction Tasks
- Natural SQL: Making SQL Easier to Infer from Natural Language Specifications
- LAGr: Labeling Aligned Graphs for Improving Systematic Generalization in Semantic Parsing
- Grounded Graph Decoding Improves Compositional Generalization in Question Answering
- Learning to Generalize Compositionally by Transferring Across Semantic Parsing Tasks