Pre-trained Models for Natural Language Processing: A Survey
arXiv:2003.08271 · doi:10.1007/s11431-020-1647-3
Abstract
Recently, the emergence of pre-trained models (PTMs) has brought natural language processing (NLP) to a new era. In this survey, we provide a comprehensive review of PTMs for NLP. We first briefly introduce language representation learning and its research progress. Then we systematically categorize existing PTMs based on a taxonomy with four perspectives. Next, we describe how to adapt the knowledge of PTMs to the downstream tasks. Finally, we outline some potential directions of PTMs for future research. This survey is purposed to be a hands-on guide for understanding, using, and developing PTMs for various NLP tasks.
Invited Review of Science China Technological Sciences
References in corpus (72)
- Distilling the Knowledge in a Neural Network
- Sequence to Sequence Learning with Neural Networks
- Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling
- Semi-Supervised Classification with Graph Convolutional Networks
- Natural Language Processing (almost) from Scratch
- Distributed Representations of Sentences and Documents
- Neural Architecture Search with Reinforcement Learning
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Cross-lingual Language Model Pretraining
- Pre-trained Models for Natural Language Processing: A Survey
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- ERNIE: Enhanced Representation through Knowledge Integration
- CamemBERT: a Tasty French Language Model
- Publicly Available Clinical BERT Embeddings
- Skip-Thought Vectors
- AraBERT: Transformer-based Model for Arabic Language Understanding
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers
- Multilingual Denoising Pre-training for Neural Machine Translation
- MASS: Masked Sequence to Sequence Pre-training for Language Generation
- Attention is not Explanation
- Q8BERT: Quantized 8Bit BERT
- Well-Read Students Learn Better: On the Importance of Pre-training Compact Models
- BERT Post-Training for Review Reading Comprehension and Aspect-based Sentiment Analysis
- Utilizing BERT for Aspect-Based Sentiment Analysis via Constructing Auxiliary Sentence
- Distilling Task-Specific Knowledge from BERT into Simple Neural Networks
- Learning How to Ask: Querying LMs with Mixtures of Soft Prompts
- Prefix-Tuning: Optimizing Continuous Prompts for Generation
- Assessing BERT's Syntactic Abilities
- Adaptation of Deep Bidirectional Multilingual Transformers for Russian Language
- UniLMv2: Pseudo-Masked Language Models for Unified Language Model Pre-Training
- BERTje: A Dutch BERT Model
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- A Theoretical Analysis of Contrastive Unsupervised Representation Learning
- Incorporating BERT into Neural Machine Translation
- Cross-Lingual Ability of Multilingual BERT: An Empirical Study
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Parameter-Efficient Transfer Learning for NLP
- What do you learn from context? Probing for sentence structure in contextualized word representations
- Data Augmentation using Pre-trained Transformer Models
- K-Adapter: Infusing Knowledge into Pre-Trained Models with Adapters
- Making Pre-trained Language Models Better Few-shot Learners
- Multilingual is not enough: BERT for Finnish
- BERT and PALs: Projected Attention Layers for Efficient Adaptation in Multi-Task Learning
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- MobileBERT: a Compact Task-Agnostic BERT for Resource-Limited Devices
- Adversarial Training for Large Neural Language Models
- PatentBERT: Patent Classification with Fine-Tuning a pre-trained BERT Model
- A Mutual Information Maximization Perspective of Language Representation Learning
- KEPLER: A Unified Model for Knowledge Embedding and Pre-trained Language Representation
- Adv-BERT: BERT is not robust on misspellings! Generating nature adversarial samples on BERT
- BERT-ATTACK: Adversarial Attack Against BERT Using BERT
- FlauBERT: Unsupervised Language Model Pre-training for French
- Utilizing BERT Intermediate Layers for Aspect Based Sentiment Analysis and Natural Language Inference
- FastBERT: a Self-distilling BERT with Adaptive Inference Time
- Exploiting BERT for End-to-End Aspect-based Sentiment Analysis
- exBERT: A Visual Analysis Tool to Explore Learned Representations in Transformers Models
- BERT Loses Patience: Fast and Robust Inference with Early Exit
- Technical report on Conversational Question Answering
- Integrating Graph Contextualized Knowledge into Pre-trained Language Models
- Are Pre-trained Language Models Aware of Phrases? Simple but Strong Baselines for Grammar Induction
- BERT-of-Theseus: Compressing BERT by Progressive Module Replacing
- Progress Notes Classification and Keyword Extraction using Attention-based Deep Learning Models with BERT
- TextBrewer: An Open-Source Knowledge Distillation Toolkit for Natural Language Processing
- SpeechBERT: An Audio-and-text Jointly Learned Language Model for End-to-end Spoken Question Answering
- TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
- Pretrained Encyclopedia: Weakly Supervised Knowledge-Pretrained Language Model
- Factual Probing Is [MASK]: Learning vs. Learning to Recall
- Transformer to CNN: Label-scarce distillation for efficient text classification
- A Knowledge-Enhanced Pretraining Model for Commonsense Story Generation
- Early Exiting with Ensemble Internal Classifiers
- ZEN: Pre-training Chinese Text Encoder Enhanced by N-gram Representations
Cited by in corpus (132)
- A Survey on Visual Transformer
- Pre-trained Models for Natural Language Processing: A Survey
- Unifying Large Language Models and Knowledge Graphs: A Roadmap
- Self-supervised Learning: Generative or Contrastive
- Hands-on Bayesian Neural Networks -- a Tutorial for Deep Learning Users
- A Survey of Human-in-the-loop for Machine Learning
- FedKD: Communication Efficient Federated Learning via Knowledge Distillation
- EEG based Emotion Recognition: A Tutorial and Review
- Deep Learning Based Text Classification: A Comprehensive Review
- Self-Supervised Speech Representation Learning: A Review
- Pre-training Enhanced Spatial-temporal Graph Neural Network for Multivariate Time Series Forecasting
- Towards a Unified View of Parameter-Efficient Transfer Learning
- VLP: A Survey on Vision-Language Pre-training
- A Survey of Knowledge-Enhanced Text Generation
- Natural Language Processing for Smart Healthcare
- Controllable Protein Design with Language Models
- Graph Neural Networks: Taxonomy, Advances and Trends
- PalmTree: Learning an Assembly Language Model for Instruction Embedding
- Compressing Large-Scale Transformer-Based Models: A Case Study on BERT
- ConvBERT: Improving BERT with Span-based Dynamic Convolution
- Machine Knowledge: Creation and Curation of Comprehensive Knowledge Bases
- CPT: A Pre-Trained Unbalanced Transformer for Both Chinese Language Understanding and Generation
- The Impact of Multiple Parallel Phrase Suggestions on Email Input and Composition Behaviour of Native and Non-Native English Writers
- Deep Learning for Android Malware Defenses: a Systematic Literature Review
- Paradigm Shift in Natural Language Processing
- Does the Magic of BERT Apply to Medical Code Assignment? A Quantitative Study
- Cluster-Level Contrastive Learning for Emotion Recognition in Conversations
- Backdoor Pre-trained Models Can Transfer to All
- Linguistically inspired roadmap for building biologically reliable protein language models
- LC-LLM: Explainable Lane-Change Intention and Trajectory Predictions with Large Language Models
- A Review of Predictive and Contrastive Self-supervised Learning for Medical Images
- A Survey on Multilingual Large Language Models: Corpora, Alignment, and Bias
- CTRAN: CNN-Transformer-based Network for Natural Language Understanding
- Mask-guided BERT for Few Shot Text Classification
- FoodSAM: Any Food Segmentation
- PTM4Tag: Sharpening Tag Recommendation of Stack Overflow Posts with Pre-trained Models
- Model Pruning Enables Localized and Efficient Federated Learning for Yield Forecasting and Data Sharing
- Distributed Deep Reinforcement Learning: A Survey and A Multi-Player Multi-Agent Learning Toolbox
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- GANSlider: How Users Control Generative Models for Images using Multiple Sliders with and without Feedforward Information
- Achieving Peak Performance for Large Language Models: A Systematic Review
- Natural Language Interfaces to Data
- AI in Human-computer Gaming: Techniques, Challenges and Opportunities
- Automated scholarly paper review: Concepts, technologies, and challenges
- Pre-trained Language Models in Biomedical Domain: A Systematic Survey
- Lightweight Transformers for Clinical Natural Language Processing
- Are Pre-trained Convolutions Better than Pre-trained Transformers?
- LogPrécis: Unleashing Language Models for Automated Malicious Log Analysis
- Graph Foundation Models: Concepts, Opportunities and Challenges
- Gender Bias in BERT -- Measuring and Analysing Biases through Sentiment Rating in a Realistic Downstream Classification Task
- PTR: Prompt Tuning with Rules for Text Classification
- A Benchmark for Lease Contract Review
- Training-free Lexical Backdoor Attacks on Language Models
- Knowledge Enhanced Pretrained Language Models: A Compreshensive Survey
- A Survey of Knowledge Enhanced Pre-trained Models
- Mask and Cloze: Automatic Open Cloze Question Generation using a Masked Language Model
- Disentangled Variational Autoencoder for Emotion Recognition in Conversations
- Unstructured Text Enhanced Open-domain Dialogue System: A Systematic Survey
- CoLAKE: Contextualized Language and Knowledge Embedding
- More but Correct: Generating Diversified and Entity-revised Medical Response
- Is Your Model Sensitive? SPeDaC: A New Benchmark for Detecting and Classifying Sensitive Personal Data
- Artificial intelligence is algorithmic mimicry: why artificial "agents" are not (and won't be) proper agents
- Pretrained Language Models for Text Generation: A Survey
- Early Exiting with Ensemble Internal Classifiers
- The Life Cycle of Knowledge in Big Language Models: A Survey
- Towards an Enhanced Understanding of Bias in Pre-trained Neural Language Models: A Survey with Special Emphasis on Affective Bias
- NodePiece: Compositional and Parameter-Efficient Representations of Large Knowledge Graphs
- A Multimodal Late Fusion Model for E-Commerce Product Classification
- Explainable Recommendation with Personalized Review Retrieval and Aspect Learning
- Ranking and Tuning Pre-trained Models: A New Paradigm for Exploiting Model Hubs
- Morphosyntactic probing of multilingual BERT models
- Multimodal Side-Tuning for Document Classification
- A Unified Generative Framework for Aspect-Based Sentiment Analysis
- PTUM: Pre-training User Model from Unlabeled User Behaviors via Self-supervision
- On the Effectiveness of Pretrained Models for API Learning
- SCARF: Self-Supervised Contrastive Learning using Random Feature Corruption
- Infusing Disease Knowledge into BERT for Health Question Answering, Medical Inference and Disease Name Recognition
- Emergent Language: A Survey and Taxonomy
- Literature review on vulnerability detection using NLP technology
- Learning Bidirectional Translation between Descriptions and Actions with Small Paired Data
- Uni-Perceiver: Pre-training Unified Architecture for Generic Perception for Zero-shot and Few-shot Tasks
- Scattered or Connected? An Optimized Parameter-efficient Tuning Approach for Information Retrieval
- Self-Supervised Learning for Text Recognition: A Critical Survey
- Encoder-Decoder Models Can Benefit from Pre-trained Masked Language Models in Grammatical Error Correction
- Accurate, yet inconsistent? Consistency Analysis on Language Understanding Models
- Exploring Motivations for Algorithm Mention in the Domain of Natural Language Processing: A Deep Learning Approach
- Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
- Enhancing Function Name Prediction using Votes-Based Name Tokenization and Multi-Task Learning
- Exploring the Innovation Opportunities for Pre-trained Models
- A Unified Generative Framework for Various NER Subtasks
- Machine Learning for Multimodal Electronic Health Records-based Research: Challenges and Perspectives
- A Large-Scale Chinese Short-Text Conversation Dataset
- LRC-BERT: Latent-representation Contrastive Knowledge Distillation for Natural Language Understanding
- Empowering News Recommendation with Pre-trained Language Models
- Enhancing Source Code Classification Effectiveness via Prompt Learning Incorporating Knowledge Features
- Accelerating BERT Inference for Sequence Labeling via Early-Exit
- Attentive Representation Learning with Adversarial Training for Short Text Clustering
- One Teacher is Enough? Pre-trained Language Model Distillation from Multiple Teachers
- Low-Resource Multi-Granularity Academic Function Recognition Based on Multiple Prompt Knowledge
- Transfer Learning for Multi-lingual Tasks -- a Survey
- An Interdisciplinary Perspective on Evaluation and Experimental Design for Visual Text Analytics: Position Paper
- NLP is Not enough -- Contextualization of User Input in Chatbots
- Beyond Pixels: Leveraging the Language of Soccer to Improve Spatio-Temporal Action Detection in Broadcast Videos
- A Graph Representation of Semi-structured Data for Web Question Answering
- Knowledge Transfer via Pre-training for Recommendation: A Review and Prospect
- NewsBERT: Distilling Pre-trained Language Model for Intelligent News Application
- NPS-AntiClone: Identity Cloning Detection based on Non-Privacy-Sensitive User Profile Data
- Approximating Human Strategic Reasoning with LLM-Enhanced Recursive Reasoners Leveraging Multi-agent Hypergames
- LARR: Large Language Model Aided Real-time Scene Recommendation with Semantic Understanding
- CCAE: A Corpus of Chinese-based Asian Englishes
- The meaning of prompts and the prompts of meaning: Semiotic reflections and modelling
- SimCLAD: A Simple Framework for Contrastive Learning of Acronym Disambiguation
- Human Activity Recognition in an Open World
- SMedBERT: A Knowledge-Enhanced Pre-trained Language Model with Structured Semantics for Medical Text Mining
- Backdoor Attacks on Pre-trained Models by Layerwise Weight Poisoning
- fastHan: A BERT-based Multi-Task Toolkit for Chinese NLP
- TCube: Domain-Agnostic Neural Time-series Narration
- Cross-lingual Intermediate Fine-tuning improves Dialogue State Tracking
- Scalability vs. Utility: Do We Have to Sacrifice One for the Other in Data Importance Quantification?
- Want to Identify, Extract and Normalize Adverse Drug Reactions in Tweets? Use RoBERTa
- Weight Squeezing: Reparameterization for Knowledge Transfer and Model Compression
- Discriminative, Generative and Self-Supervised Approaches for Target-Agnostic Learning
- Pre-training and Diagnosing Knowledge Base Completion Models
- Navigating the Kaleidoscope of COVID-19 Misinformation Using Deep Learning
- Efficient Attribute Injection for Pretrained Language Models
- Multilingual Translation via Grafting Pre-trained Language Models
- Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap
- Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
- A Survey on Green Deep Learning
- Position-based Contributive Embeddings for Aspect-Based Sentiment Analysis
- REPT: Bridging Language Models and Machine Reading Comprehension via Retrieval-Based Pre-training
- Event-Driven Learning of Systematic Behaviours in Stock Markets