Measuring Compositional Generalization: A Comprehensive Method on Realistic Data
arXiv:1912.09713
Abstract
State-of-the-art machine learning methods exhibit limited compositional generalization. At the same time, there is a lack of realistic benchmarks that comprehensively measure this ability, which makes it challenging to find and evaluate improvements. We introduce a novel method to systematically construct such benchmarks by maximizing compound divergence while guaranteeing a small atom divergence between train and test sets, and we quantitatively compare this method to other approaches for creating compositional generalization benchmarks. We present a large and realistic natural language question answering dataset that is constructed according to this method, and we use it to analyze the compositional generalization ability of three machine learning architectures. We find that they fail to generalize compositionally and that there is a surprisingly strong negative correlation between compound divergence and accuracy. We also demonstrate how our method can be used to create new compositionality benchmarks on top of the existing SCAN dataset, which confirms these findings.
Accepted for publication at ICLR 2020
References in corpus (5)
- Improving Text-to-SQL Evaluation Methodology
- Variational Reasoning for Question Answering with Knowledge Graph
- Compositional generalization in a deep seq2seq model by separating syntax and semantics
- SCAN: Learning Hierarchical Compositional Visual Concepts
- Automatically Composing Representation Transformations as a Means for Generalization
Cited by in corpus (25)
- A Survey of Zero-shot Generalisation in Deep Reinforcement Learning
- Compositional Generalization in Semantic Parsing: Pre-training vs. Specialized Architectures
- A Benchmark for Systematic Generalization in Grounded Language Understanding
- Language as a Cognitive Tool to Imagine Goals in Curiosity-Driven Exploration
- On the Binding Problem in Artificial Neural Networks
- CLOSURE: Assessing Systematic Generalization of CLEVR Models
- Compositional Generalization by Learning Analytical Expressions
- Measuring Systematic Generalization in Neural Proof Generation with Transformers
- The Scattering Compositional Learner: Discovering Objects, Attributes, Relationships in Analogical Reasoning
- Hierarchical Poset Decoding for Compositional Generalization in Language
- RuBQ: A Russian Dataset for Question Answering over Wikidata
- Learning to Recombine and Resample Data for Compositional Generalization
- Dynamic Inference with Neural Interpreters
- Modularity in Deep Learning: A Survey
- Compositional Generalization via Neural-Symbolic Stack Machines
- Compositional Generalization via Semantic Tagging
- Revisiting Iterative Back-Translation from the Perspective of Compositional Generalization
- Pointer Value Retrieval: A new benchmark for understanding the limits of neural network generalization
- Automatic Knowledge Augmentation for Generative Commonsense Reasoning
- SQALER: Scaling Question Answering by Decoupling Multi-Hop and Logical Reasoning
- LAGr: Labeling Aligned Graphs for Improving Systematic Generalization in Semantic Parsing
- Revisit Systematic Generalization via Meaningful Learning
- Learning to Generalize Compositionally by Transferring Across Semantic Parsing Tasks
- How BPE Affects Memorization in Transformers
- Complex Knowledge Base Question Answering: A Survey