ERNIE: Enhanced Representation through Knowledge Integration
arXiv:1904.09223
Abstract
We present a novel language representation model enhanced by knowledge called ERNIE (Enhanced Representation through kNowledge IntEgration). Inspired by the masking strategy of BERT, ERNIE is designed to learn language representation enhanced by knowledge masking strategies, which includes entity-level masking and phrase-level masking. Entity-level strategy masks entities which are usually composed of multiple words.Phrase-level strategy masks the whole phrase which is composed of several words standing together as a conceptual unit.Experimental results show that ERNIE outperforms other baseline methods, achieving new state-of-the-art results on five Chinese natural language processing tasks including natural language inference, semantic similarity, named entity recognition, sentiment analysis and question answering. We also demonstrate that ERNIE has more powerful knowledge inference capacity on a cloze test.
8 pages
References in corpus (3)
Cited by in corpus (20)
- ERNIE: Enhanced Language Representation with Informative Entities
- KERMIT: Generative Insertion-Based Modeling for Sequences
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data
- SKEP: Sentiment Knowledge Enhanced Pre-training for Sentiment Analysis
- A Pre-training Based Personalized Dialogue Generation Model with Persona-sparse Data
- A Survey on Self-supervised Pre-training for Sequential Transfer Learning in Neural Networks
- A Knowledge-Enhanced Pretraining Model for Commonsense Story Generation
- To Tune or Not To Tune? How About the Best of Both Worlds?
- PMI-Masking: Principled masking of correlated spans
- ZEN: Pre-training Chinese Text Encoder Enhanced by N-gram Representations
- Synthetic QA Corpora Generation with Roundtrip Consistency
- Hardware Accelerator for Multi-Head Attention and Position-Wise Feed-Forward in the Transformer
- Accenture at CheckThat! 2020: If you say so: Post-hoc fact-checking of claims using transformer-based models
- ERNIE at SemEval-2020 Task 10: Learning Word Emphasis Selection by Pre-trained Language Model
- Which *BERT? A Survey Organizing Contextualized Encoders
- Application of Pre-training Models in Named Entity Recognition
- MATINF: A Jointly Labeled Large-Scale Dataset for Classification, Question Answering and Summarization
- Retrieve Synonymous keywords for Frequent Queries in Sponsored Search in a Data Augmentation Way
- Neural Mask Generator: Learning to Generate Adaptive Word Maskings for Language Model Adaptation
- BIG MOOD: Relating Transformers to Explicit Commonsense Knowledge