Representation of linguistic form and function in recurrent neural networks
arXiv:1602.08952
Abstract
We present novel methods for analyzing the activation patterns of RNNs from a linguistic point of view and explore the types of linguistic structure they learn. As a case study, we use a multi-task gated recurrent network architecture consisting of two parallel pathways with shared word embeddings trained on predicting the representations of the visual scene corresponding to an input sentence, and predicting the next word in the same sentence. Based on our proposed method to estimate the amount of contribution of individual tokens in the input to the final prediction of the networks we show that the image prediction pathway: a) is sensitive to the information structure of the sentence b) pays selective attention to lexical categories and grammatical functions that carry semantic information c) learns to treat the same input token differently depending on its grammatical functions in the sentence. In contrast the language model is comparatively more sensitive to words with a syntactic function. Furthermore, we propose methods to ex- plore the function of individual hidden units in RNNs and show that the two pathways of the architecture in our case study contain specialized units tuned to patterns informative for the task, some of which can carry activations to later time steps to encode long-term dependencies.
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Understanding Neural Networks Through Deep Visualization
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- DRAW: A Recurrent Neural Network For Image Generation
- Multifaceted Feature Visualization: Uncovering the Different Types of Features Learned By Each Neuron in Deep Neural Networks
- Learning language through pictures
Cited by in corpus (13)
- Understanding Neural Networks through Representation Erasure
- Fine-grained Analysis of Sentence Embeddings Using Auxiliary Prediction Tasks
- What do Neural Machine Translation Models Learn about Morphology?
- Imagination improves Multimodal Translation
- Evaluating Layers of Representation in Neural Machine Translation on Part-of-Speech and Semantic Tagging Tasks
- Identifying and Controlling Important Neurons in Neural Machine Translation
- Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies
- From phonemes to images: levels of representation in a recurrent neural model of visually-grounded language learning
- LSTMVis: A Tool for Visual Analysis of Hidden State Dynamics in Recurrent Neural Networks
- Memory Visualization for Gated Recurrent Neural Networks in Speech Recognition
- Indicatements that character language models learn English morpho-syntactic units and regularities
- Investigating how well contextual features are captured by bi-directional recurrent neural network models
- Sparsity Emerges Naturally in Neural Language Models