Jointly Learning Sentence Embeddings and Syntax with Unsupervised Tree-LSTMs
arXiv:1705.09189 · doi:10.1017/S1351324919000184
Abstract
We introduce a neural network that represents sentences by composing their words according to induced binary parse trees. We use Tree-LSTM as our composition function, applied along a tree structure found by a fully differentiable natural language chart parser. Our model simultaneously optimises both the composition function and the parser, thus eliminating the need for externally-provided parse trees which are normally required for Tree-LSTM. It can therefore be seen as a tree-based RNN that is unsupervised with respect to the parse trees. As it is fully differentiable, our model is easily trained with an off-the-shelf gradient descent method and backpropagation. We demonstrate that it achieves better performance compared to various supervised Tree-LSTM architectures on a textual entailment task and a reverse dictionary task.
References in corpus (9)
- Sequence to Sequence Learning with Neural Networks
- Overcoming catastrophic forgetting in neural networks
- Categorical Reparameterization with Gumbel-Softmax
- Layer Normalization
- DyNet: The Dynamic Neural Network Toolkit
- DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- Learning to Compose Words into Sentences with Reinforcement Learning
- Latent Tree Learning with Differentiable Parsers: Shift-Reduce Parsing and Chart Parsing
Cited by in corpus (28)
- Learning to Compose and Reason with Language Tree Structures for Visual Grounding
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
- GLoMo: Unsupervisedly Learned Relational Graphs as Transferable Representations
- Unsupervised Latent Tree Induction with Deep Inside-Outside Recursive Autoencoders
- Latent Tree Learning with Differentiable Parsers: Shift-Reduce Parsing and Chart Parsing
- "When they say weed causes depression, but it's your fav antidepressant": Knowledge-aware Attention Framework for Relationship Extraction
- R2D2: Recursive Transformer based on Differentiable Tree for Interpretable Hierarchical Language Modeling
- BERE: An accurate distantly supervised biomedical entity relation extraction network
- "Is depression related to cannabis?": A knowledge-infused model for Entity and Relation Extraction with Limited Supervision
- Modeling Latent Sentence Structure in Neural Machine Translation
- Tensor Product Generation Networks for Deep NLP Modeling
- Latent Compositional Representations Improve Systematic Generalization in Grounded Question Answering
- Weakly Supervised Reasoning by Neuro-Symbolic Approaches
- Cooperative Learning of Disjoint Syntax and Semantics
- Sentence Encoding with Tree-constrained Relation Networks
- An Imitation Learning Approach to Unsupervised Parsing
- Neural Machine Translation: A Review and Survey
- Dialogue Act Classification in Group Chats with DAG-LSTMs
- Dynamic Compositionality in Recursive Neural Networks with Structure-aware Tag Representations
- Learning to Embed Sentences Using Attentive Recursive Trees
- On Tree-Based Neural Sentence Modeling
- Neural Compositional Denotational Semantics for Question Answering
- On learning an interpreted language with recurrent models
- Latent Tree Learning with Ordered Neurons: What Parses Does It Produce?
- Should Semantic Vector Composition be Explicit? Can it be Linear?
- To be Closer: Learning to Link up Aspects with Opinions
- Attentive Tensor Product Learning
- Towards Dynamic Computation Graphs via Sparse Latent Structure