Multi-Task Cross-Lingual Sequence Tagging from Scratch
arXiv:1603.06270
Abstract
We present a deep hierarchical recurrent neural network for sequence tagging. Given a sequence of words, our model employs deep gated recurrent units on both character and word levels to encode morphology and context information, and applies a conditional random field layer to predict the tags. Our model is task independent, language independent, and feature engineering free. We further extend our model to multi-task and cross-lingual joint training by sharing the architecture and parameters. Our model achieves state-of-the-art results in multiple languages on several benchmark tasks including POS tagging, chunking, and NER. We also demonstrate that multi-task and cross-lingual joint training can improve the performance in various cases.
References in corpus (5)
- Sequence to Sequence Learning with Neural Networks
- Natural Language Processing (almost) from Scratch
- On the Properties of Neural Machine Translation: Encoder-Decoder Approaches
- Finding Function in Form: Compositional Character Models for Open Vocabulary Word Representation
- Multilingual Language Processing From Bytes
Cited by in corpus (46)
- A Survey on Deep Learning for Named Entity Recognition
- A Survey on Recent Advances in Named Entity Recognition from Deep Learning models
- Revisiting Semi-Supervised Learning with Graph Embeddings
- End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF
- On the Origin of Deep Learning
- Semantic Tagging with Deep Residual Networks
- Learning by Association - A versatile semi-supervised training method for neural networks
- Modeling Noisiness to Recognize Named Entities using Multitask Neural Networks on Social Media
- Words or Characters? Fine-grained Gating for Reading Comprehension
- Joint Entity Extraction and Assertion Detection for Clinical Text
- Neural Models for Sequence Chunking
- Deep Active Learning for Named Entity Recognition
- DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks
- Enhance word representation for out-of-vocabulary on Ubuntu dialogue corpus
- Named Entity Recognition with stack residual LSTM and trainable bias decoding
- Subword-augmented Embedding for Cloze Reading Comprehension
- Improving Named Entity Recognition by Jointly Learning to Disambiguate Morphological Tags
- A Survey on Recent Advances in Sequence Labeling from Deep Learning Models
- Semi-Supervised Disentangled Framework for Transferable Named Entity Recognition
- Enhancing deep neural networks with morphological information
- Hierarchical Contextualized Representation for Named Entity Recognition
- Sequence Labeling: A Practical Approach
- Neural Cross-Lingual Named Entity Recognition with Minimal Resources
- GCDT: A Global Context Enhanced Deep Transition Architecture for Sequence Labeling
- Morphological Embeddings for Named Entity Recognition in Morphologically Rich Languages
- Byte-based Language Identification with Deep Convolutional Networks
- Meta-Learning Multi-task Communication
- On Difficulties of Cross-Lingual Transfer with Order Differences: A Case Study on Dependency Parsing
- Deep Semi-Supervised Learning with Linguistically Motivated Sequence Labeling Task Hierarchies
- Joint Multi-Domain Learning for Automatic Short Answer Grading
- Multi-task Learning over Graph Structures
- Chinese Discourse Segmentation Using Bilingual Discourse Commonality
- Neural Named Entity Recognition from Subword Units
- Does Higher Order LSTM Have Better Accuracy for Segmenting and Labeling Sequence Data?
- Multi-task Domain Adaptation for Sequence Tagging
- Improving Aspect-Level Sentiment Analysis with Aspect Extraction
- Effective Subword Segmentation for Text Comprehension
- Domain Adaptation for Neural Networks by Parameter Augmentation
- KnowNER: Incremental Multilingual Knowledge in Named Entity Recognition
- A Multi-task Learning Approach for Named Entity Recognition using Local Detection
- Scaling Matters in Deep Structured-Prediction Models
- Empirical Evaluation of Multi-task Learning in Deep Neural Networks for Natural Language Processing
- Converse Attention Knowledge Transfer for Low-Resource Named Entity Recognition
- Multi-Task Learning for Argumentation Mining
- Multi-task Learning for Chinese Word Usage Errors Detection
- Effective Context and Fragment Feature Usage for Named Entity Recognition