TENER: Adapting Transformer Encoder for Named Entity Recognition
arXiv:1911.04474
Abstract
The Bidirectional long short-term memory networks (BiLSTM) have been widely used as an encoder in models solving the named entity recognition (NER) task. Recently, the Transformer is broadly adopted in various Natural Language Processing (NLP) tasks owing to its parallelism and advantageous performance. Nevertheless, the performance of the Transformer in NER is not as good as it is in other NLP tasks. In this paper, we propose TENER, a NER architecture adopting adapted Transformer Encoder to model the character-level features and word-level features. By incorporating the direction and relative distance aware attention and the un-scaled attention, we prove the Transformer-like encoder is just as effective for NER as other NLP tasks.
Corrept typos, update performance based on the public available codes
References in corpus (2)
Cited by in corpus (13)
- A Survey on Recent Advances in Sequence Labeling from Deep Learning Models
- FLAT: Chinese NER Using Flat-Lattice Transformer
- Simplified DOM Trees for Transferable Attribute Extraction from the Web
- A Unified Generative Framework for Various NER Subtasks
- Improving Named Entity Recognition with Attentive Ensemble of Syntactic Information
- Accelerating BERT Inference for Sequence Labeling via Early-Exit
- PiSLTRc: Position-informed Sign Language Transformer with Content-aware Convolution
- Named Entity Recognition for Social Media Texts with Semantic Augmentation
- BERT for Monolingual and Cross-Lingual Reverse Dictionary
- Multi-Scale Local-Temporal Similarity Fusion for Continuous Sign Language Recognition
- Larger-Context Tagging: When and Why Does It Work?
- A Partition Filter Network for Joint Entity and Relation Extraction
- Cascaded Semantic and Positional Self-Attention Network for Document Classification