When Are Tree Structures Necessary for Deep Learning of Representations?
arXiv:1503.00185
Abstract
Recursive neural models, which use syntactic parse trees to recursively generate representations bottom-up, are a popular architecture. But there have not been rigorous evaluations showing for exactly which tasks this syntax-based method is appropriate. In this paper we benchmark {\bf recursive} neural models against sequential {\bf recurrent} neural models (simple recurrent and LSTM models), enforcing apples-to-apples comparison as much as possible. We investigate 4 tasks: (1) sentiment classification at the sentence level and phrase level; (2) matching questions to answer-phrases; (3) discourse parsing; (4) semantic relation extraction (e.g., {\em component-whole} between nouns). Our goal is to understand better when, and why, recursive models can outperform simpler models. We find that recursive models help mainly on tasks (like semantic relation extraction) that require associating headwords across a long distance, particularly on very long sequences. We then introduce a method for allowing recurrent models to achieve similar performance: breaking long sentences into clause-like units at punctuation and processing them separately before combining. Our results thus help understand the limitations of both classes of models, and suggest directions for improving recurrent models.
References in corpus (5)
Cited by in corpus (41)
- A C-LSTM Neural Network for Text Classification
- End-to-End Relation Extraction using LSTMs on Sequences and Tree Structures
- Glyce: Glyph-vectors for Chinese Character Representations
- Neural Machine Translation and Sequence-to-sequence Models: A Tutorial
- DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding
- Aspect Level Sentiment Classification with Deep Memory Network
- Combining Neural Networks and Log-linear Models to Improve Relation Extraction
- Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation
- BP-Transformer: Modelling Long-Range Context via Binary Partitioning
- Improved Neural Machine Translation with a Syntax-Aware Encoder and Decoder
- A Fast Unified Model for Parsing and Sentence Understanding
- Description Based Text Classification with Reinforcement Learning
- DRAGNN: A Transition-based Framework for Dynamically Connected Neural Networks
- Bidirectional Tree-Structured LSTM with Head Lexicalization
- Interpreting Deep Learning Models in Natural Language Processing: A Review
- Hierarchical Character Embeddings: Learning Phonological and Semantic Representations in Languages of Logographic Origin using Recursive Neural Networks
- SEE: Syntax-aware Entity Embedding for Neural Relation Extraction
- Combining Convolution and Recursive Neural Networks for Sentiment Analysis
- What Do Recurrent Neural Network Grammars Learn About Syntax?
- Neural Network Models for Implicit Discourse Relation Classification in English and Chinese without Surface Features
- Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability
- Full-Time Supervision based Bidirectional RNN for Factoid Question Answering
- Equation Embeddings
- Multi-Scale Self-Attention for Text Classification
- T3: Tree-Autoencoder Constrained Adversarial Text Generation for Targeted Attack
- Stochastic Learning of Nonstationary Kernels for Natural Language Modeling
- Dialogue History Matters! Personalized Response Selectionin Multi-turn Retrieval-based Chatbots
- Tag-Enhanced Tree-Structured Neural Networks for Implicit Discourse Relation Classification
- Neural Machine Translation with Source-Side Latent Graph Parsing
- A Framework for End-to-End Learning on Semantic Tree-Structured Data
- Universal Dependencies Parsing for Colloquial Singaporean English
- Syntax-Enhanced Neural Machine Translation with Syntax-Aware Word Representations
- Shallow Discourse Parsing Using Distributed Argument Representations and Bayesian Optimization
- On Tree-Based Neural Sentence Modeling
- New and Improved Algorithms for Unordered Tree Inclusion
- A constrained recursion algorithm for batch normalization of tree-sturctured LSTM
- What You Say and How You Say it: Joint Modeling of Topics and Discourse in Microblog Conversations
- Latent Variable Sentiment Grammar
- Multiple Structural Priors Guided Self Attention Network for Language Understanding
- A More Efficient Chinese Named Entity Recognition base on BERT and Syntactic Analysis
- Improving Neural Sequence Labelling using Additional Linguistic Information