The Importance of Being Recurrent for Modeling Hierarchical Structure
arXiv:1803.03585
Abstract
Recent work has shown that recurrent neural networks (RNNs) can implicitly capture and exploit hierarchical information when trained to solve common natural language processing tasks such as language modeling (Linzen et al., 2016) and neural machine translation (Shi et al., 2016). In contrast, the ability to model structured data with non-recurrent neural networks has received little attention despite their success in many NLP tasks (Gehring et al., 2017; Vaswani et al., 2017). In this work, we compare the two architectures---recurrent versus non-recurrent---with respect to their ability to model hierarchical structure and find that recurrency is indeed important for this purpose.
EMNLP 2018
References in corpus (4)
Cited by in corpus (6)
- Why Self-Attention? A Targeted Evaluation of Neural Machine Translation Architectures
- Addressing Some Limitations of Transformers with Feedback Memory
- Compositionality decomposed: how do neural networks generalise?
- What can linguistics and deep learning contribute to each other?
- Understanding Cross-Lingual Syntactic Transfer in Multilingual Recurrent Neural Networks
- SimpleBooks: Long-term dependency book dataset with simplified English vocabulary for word-level language modeling