Neural Probabilistic Model for Non-projective MST Parsing
arXiv:1701.00874
Abstract
In this paper, we propose a probabilistic parsing model, which defines a proper conditional probability distribution over non-projective dependency trees for a given sentence, using neural representations as inputs. The neural network architecture is based on bi-directional LSTM-CNNs which benefits from both word- and character-level representations automatically, by using combination of bidirectional LSTM and CNN. On top of the neural network, we introduce a probabilistic structured layer, defining a conditional log-linear model over non-projective trees. We evaluate our model on 17 different datasets, across 14 different languages. By exploiting Kirchhoff's Matrix-Tree Theorem (Tutte, 1984), the partition functions and marginals can be computed efficiently, leading to a straight-forward end-to-end model training procedure via back-propagation. Our parser achieves state-of-the-art parsing performance on nine datasets.
To appear in IJCNLP 2017
References in corpus (10)
- ADADELTA: An Adaptive Learning Rate Method
- Natural Language Processing (almost) from Scratch
- Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
- Transition-Based Dependency Parsing with Stack Long Short-Term Memory
- Simple and Accurate Dependency Parsing Using Bidirectional LSTM Feature Representations
- Training with Exploration Improves a Greedy Stack-LSTM Parser
- Bi-directional Attention with Agreement for Dependency Parsing
- Dropout with Expectation-linear Regularization
- Probabilistic Models for High-Order Projective Dependency Parsing
- Dependency Parsing as Head Selection
Cited by in corpus (10)
- Towards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation
- Deep Multitask Learning for Semantic Dependency Parsing
- Head-Driven Phrase Structure Grammar Parsing on Penn Treebank
- Semi-Supervised Sequence Modeling with Cross-View Training
- Second-Order Neural Dependency Parsing with Message Passing and End-to-End Training
- An improved neural network model for joint POS tagging and dependency parsing
- Two Local Models for Neural Constituent Parsing
- Multitask Pointer Network for Multi-Representational Parsing
- Sentence Encoding with Tree-constrained Relation Networks
- Efficient Sampling of Dependency Structures