Sentence Encoders on STILTs: Supplementary Training on Intermediate Labeled-data Tasks
arXiv:1811.01088
Abstract
Pretraining sentence encoders with language modeling and related unsupervised tasks has recently been shown to be very effective for language understanding tasks. By supplementing language model-style pretraining with further training on data-rich supervised tasks, such as natural language inference, we obtain additional performance improvements on the GLUE benchmark. Applying supplementary training on BERT (Devlin et al., 2018), we attain a GLUE score of 81.8---the state of the art (as of 02/24/2019) and a 1.4 point improvement over BERT. We also observe reduced variance across random restarts in this setting. Our approach yields similar improvements when applied to ELMo (Peters et al., 2018a) and Radford et al. (2018)'s model. In addition, the benefits of supplementary training are particularly pronounced in data-constrained regimes, as we show in experiments with artificially limited training data.
References in corpus (3)
- Aligning Books and Movies: Towards Story-like Visual Explanations by Watching Movies and Reading Books
- What you can cram into a single vector: Probing sentence embeddings for linguistic properties
- Language Modeling Teaches You More Syntax than Translation Does: Lessons Learned Through Auxiliary Task Analysis
Cited by in corpus (90)
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer
- Language Models are Few-Shot Learners
- Pre-trained Models for Natural Language Processing: A Survey
- SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
- Multi-Task Deep Neural Networks for Natural Language Understanding
- Fine-Tuning Pretrained Language Models: Weight Initializations, Data Orders, and Early Stopping
- Revisiting Few-sample BERT Fine-tuning
- On the Stability of Fine-tuning BERT: Misconceptions, Explanations, and Strong Baselines
- BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions
- True Few-Shot Learning with Language Models
- Supervised Multimodal Bitransformers for Classifying Images and Text
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- Making Pre-trained Language Models Better Few-shot Learners
- Pretrained Transformers for Text Ranking: BERT and Beyond
- Entailment as Few-Shot Learner
- Mixout: Effective Regularization to Finetune Large-scale Pretrained Language Models
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- DeCLUTR: Deep Contrastive Learning for Unsupervised Textual Representations
- Do Attention Heads in BERT Track Syntactic Dependencies?
- Multi-task learning for natural language processing in the 2020s: where are we going?
- KLUE: Korean Language Understanding Evaluation
- Differentiable Prompt Makes Pre-trained Language Models Better Few-shot Learners
- Intermediate-Task Transfer Learning with Pretrained Models for Natural Language Understanding: When and Why Does It Work?
- Few-Shot Named Entity Recognition: A Comprehensive Study
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' Disagreement
- AMMUS : A Survey of Transformer-based Pretrained Models in Natural Language Processing
- Variational Information Bottleneck for Effective Low-Resource Fine-Tuning
- Transfer Fine-Tuning: A BERT Case Study
- Unifying Question Answering, Text Classification, and Regression via Span Extraction
- Pre-trained Language Model Representations for Language Generation
- jiant: A Software Toolkit for Research on General-Purpose Text Understanding Models
- Show Your Work: Improved Reporting of Experimental Results
- The MultiBERTs: BERT Reproductions for Robustness Analysis
- KorNLI and KorSTS: New Benchmark Datasets for Korean Natural Language Understanding
- Transferability of Natural Language Inference to Biomedical Question Answering
- HUBERT Untangles BERT to Improve Transfer across NLP Tasks
- On the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation
- CausalBERT: Injecting Causal Knowledge Into Pre-trained Models with Minimal Supervision
- MMM: Multi-stage Multi-task Learning for Multi-choice Reading Comprehension
- Story Ending Prediction by Transferable BERT
- Exploring and Predicting Transferability across NLP Tasks
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting
- Can You Tell Me How to Get Past Sesame Street? Sentence-Level Pretraining Beyond Language Modeling
- Improving BERT with Self-Supervised Attention
- ExT5: Towards Extreme Multi-Task Scaling for Transfer Learning
- Is Supervised Syntactic Parsing Beneficial for Language Understanding? An Empirical Investigation
- Asking Crowdworkers to Write Entailment Examples: The Best of Bad Options
- AUBER: Automated BERT Regularization
- Repulsive Attention: Rethinking Multi-head Attention as Bayesian Inference
- Task Selection Policies for Multitask Learning
- Which *BERT? A Survey Organizing Contextualized Encoders
- An Empirical Study on Hyperparameter Optimization for Fine-Tuning Pre-trained Language Models
- Train No Evil: Selective Masking for Task-Guided Pre-Training
- Domain-Relevant Embeddings for Medical Question Similarity
- Effective Transfer Learning for Identifying Similar Questions: Matching User Questions to COVID-19 FAQs
- Raise a Child in Large Language Model: Towards Effective and Generalizable Fine-tuning
- Universal Natural Language Processing with Limited Annotations: Try Few-shot Textual Entailment as a Start
- Robustness Challenges in Model Distillation and Pruning for Natural Language Understanding
- With Little Power Comes Great Responsibility
- The Curse of Performance Instability in Analysis Datasets: Consequences, Source, and Suggestions
- Human vs. Muppet: A Conservative Estimate of Human Performance on the GLUE Benchmark
- Knowledgeable Dialogue Reading Comprehension on Key Turns
- Subjective Question Answering: Deciphering the inner workings of Transformers in the realm of subjectivity
- Investigating BERT's Knowledge of Language: Five Analysis Methods with NPIs
- Robust Transfer Learning with Pretrained Language Models through Adapters
- Go Beyond Plain Fine-tuning: Improving Pretrained Models for Social Commonsense
- General Purpose Text Embeddings from Pre-trained Language Models for Scalable Inference
- A Simple and Efficient Multi-Task Learning Approach for Conditioned Dialogue Generation
- Bi-Granularity Contrastive Learning for Post-Training in Few-Shot Scene
- Convolutions and Self-Attention: Re-interpreting Relative Positions in Pre-trained Language Models
- OCHADAI-KYOTO at SemEval-2021 Task 1: Enhancing Model Generalization and Robustness for Lexical Complexity Prediction
- Few-Sample Named Entity Recognition for Security Vulnerability Reports by Fine-Tuning Pre-Trained Language Models
- The futility of STILTs for the classification of lexical borrowings in Spanish
- QuASE: Question-Answer Driven Sentence Encoding
- Uncertain Natural Language Inference
- SpartQA: : A Textual Question Answering Benchmark for Spatial Reasoning
- Learning to Generalize Compositionally by Transferring Across Semantic Parsing Tasks
- Pre-Training Transformers as Energy-Based Cloze Models
- MDQE: A More Accurate Direct Pretraining for Machine Translation Quality Estimation
- Don't Go Far Off: An Empirical Study on Neural Poetry Translation
- STraTA: Self-Training with Task Augmentation for Better Few-shot Learning
- A Closer Look at Few-Shot Crosslingual Transfer: The Choice of Shots Matters
- Overview of ADoBo 2021: Automatic Detection of Unassimilated Borrowings in the Spanish Press
- Efficient transfer learning for NLP with ELECTRA
- Cross-lingual Intermediate Fine-tuning improves Dialogue State Tracking
- Effective Unsupervised Domain Adaptation with Adversarially Trained Language Models
- Diverse Distributions of Self-Supervised Tasks for Meta-Learning in NLP
- Improved Latent Tree Induction with Distant Supervision via Span Constraints
- The Effectiveness of Intermediate-Task Training for Code-Switched Natural Language Understanding
- Beyond Fine-tuning: Few-Sample Sentence Embedding Transfer