The Devil is in the Detail: Simple Tricks Improve Systematic Generalization of Transformers
arXiv:2108.12284
Abstract
Recently, many datasets have been proposed to test the systematic generalization ability of neural networks. The companion baseline Transformers, typically trained with default hyper-parameters from standard tasks, are shown to fail dramatically. Here we demonstrate that by revisiting model configurations as basic as scaling of embeddings, early stopping, relative positional embedding, and Universal Transformer variants, we can drastically improve the performance of Transformers on systematic generalization. We report improvements on five popular datasets: SCAN, CFQ, PCFG, COGS, and Mathematics dataset. Our models improve accuracy from 50% to 85% on the PCFG productivity split, and from 35% to 81% on COGS. On SCAN, relative positional embedding largely mitigates the EOS decision problem (Newman et al., 2020), yielding 100% accuracy on the length split with a cutoff at 26. Importantly, performance differences between these models are typically invisible on the IID data split. This calls for proper generalization validation sets for developing neural networks that generalize systematically. We publicly release the code to reproduce our results.
Accepted to EMNLP 2021
References in corpus (13)
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Sequence to Sequence Learning with Neural Networks
- Sequence Transduction with Recurrent Neural Networks
- Compositional generalization in a deep seq2seq model by separating syntax and semantics
- Unlocking Compositional Generalization in Pre-trained Models Using Intermediate Representations
- On the Binding Problem in Artificial Neural Networks
- Learning Compositional Rules via Neural Program Synthesis
- Compositional Generalization by Learning Analytical Expressions
- Linear Transformers Are Secretly Fast Weight Programmers
- Self-Delimiting Neural Networks
- Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
- GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
- Learning advanced mathematical computations from examples
Cited by in corpus (5)
- Choose a Transformer: Fourier or Galerkin
- Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
- The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization
- TraVLR: Now You See It, Now You Don't! A Bimodal Dataset for Evaluating Visio-Linguistic Reasoning
- Sequence-to-Sequence Learning with Latent Neural Grammars