Improving Named Entity Recognition for Chinese Social Media with Word Segmentation Representation Learning
arXiv:1603.00786
Abstract
Named entity recognition, and other information extraction tasks, frequently use linguistic features such as part of speech tags or chunkings. For languages where word boundaries are not readily identified in text, word segmentation is a key first step to generating features for an NER system. While using word boundary tags as features are helpful, the signals that aid in identifying these boundaries may provide richer information for an NER system. New state-of-the-art word segmentation systems use neural models to learn representations for predicting word boundaries. We show that these same representations, jointly trained with an NER system, yield significant improvements in NER for Chinese social media. In our experiments, jointly training NER and word segmentation with an LSTM-CRF model yields nearly 5% absolute improvement over previously published results.
This is the camera ready version of our ACL'16 paper. We also added a supplementary material containing the results of our systems on a cleaner dataset (much higher F1 scores). More information please refer to the repo https://github.com/hltcoe/golden-horse
Cited by in corpus (10)
- End-to-end Sequence Labeling via Bi-directional LSTM-CNNs-CRF
- Cross-Sentence N-ary Relation Extraction with Graph LSTMs
- Chinese Lexical Analysis with Deep Bi-GRU-CRF Network
- Chinese NER Using Lattice LSTM
- Adversarial Transfer Learning for Punctuation Restoration
- F-Score Driven Max Margin Neural Network for Named Entity Recognition in Chinese Social Media
- Neural Adaptation Layers for Cross-domain Named Entity Recognition
- Multi-task Domain Adaptation for Sequence Tagging
- Deep Cascade Multi-task Learning for Slot Filling in Online Shopping Assistant
- Incorporating Uncertain Segmentation Information into Chinese NER for Social Media Text