Multi-Task Deep Neural Networks for Natural Language Understanding
arXiv:1901.11504
Abstract
In this paper, we present a Multi-Task Deep Neural Network (MT-DNN) for learning representations across multiple natural language understanding (NLU) tasks. MT-DNN not only leverages large amounts of cross-task data, but also benefits from a regularization effect that leads to more general representations in order to adapt to new tasks and domains. MT-DNN extends the model proposed in Liu et al. (2015) by incorporating a pre-trained bidirectional transformer language model, known as BERT (Devlin et al., 2018). MT-DNN obtains new state-of-the-art results on ten NLU tasks, including SNLI, SciTail, and eight out of nine GLUE tasks, pushing the GLUE benchmark to 82.7% (2.2% absolute improvement). We also demonstrate using the SNLI and SciTail datasets that the representations learned by MT-DNN allow domain adaptation with substantially fewer in-domain labels than the pre-trained BERT representations. The code and pre-trained models are publicly available at https://github.com/namisan/mt-dnn.
10 pages, 2 figures and 5 tables; Accepted by ACL 2019
References in corpus (2)
Cited by in corpus (28)
- ERNIE: Enhanced Representation through Knowledge Integration
- Multi-Task Learning with Deep Neural Networks: A Survey
- Improving Multi-Task Deep Neural Networks via Knowledge Distillation for Natural Language Understanding
- StructBERT: Incorporating Language Structures into Pre-training for Deep Language Understanding
- Utilizing BERT Intermediate Layers for Aspect Based Sentiment Analysis and Natural Language Inference
- MT-BioNER: Multi-task Learning for Biomedical Named Entity Recognition using Deep Bidirectional Transformers
- TwinBERT: Distilling Knowledge to Twin-Structured BERT Models for Efficient Retrieval
- To Tune or Not To Tune? How About the Best of Both Worlds?
- Towards Domain Adaptation from Limited Data for Question Answering Using Deep Neural Networks
- Multi-Task Bidirectional Transformer Representations for Irony Detection
- MMM: Multi-stage Multi-task Learning for Multi-choice Reading Comprehension
- Story Ending Prediction by Transferable BERT
- Improve Transformer Models with Better Relative Position Embeddings
- Generalization in multitask deep neural classifiers: a statistical physics approach
- Self-Supervised Dialogue Learning
- When Choosing Plausible Alternatives, Clever Hans can be Clever
- Domain-Relevant Embeddings for Medical Question Similarity
- Path-Based Contextualization of Knowledge Graphs for Textual Entailment
- Ladder Loss for Coherent Visual-Semantic Embedding
- Probing Contextualized Sentence Representations with Visual Awareness
- Enriching Conversation Context in Retrieval-based Chatbots
- Pentagon at MEDIQA 2019: Multi-task Learning for Filtering and Re-ranking Answers using Language Inference and Question Entailment
- Memeify: A Large-Scale Meme Generation System
- On the relationship between multitask neural networks and multitask Gaussian Processes
- Shareable Representations for Search Query Understanding
- Surf at MEDIQA 2019: Improving Performance of Natural Language Inference in the Clinical Domain by Adopting Pre-trained Language Model
- Pretrained Transformers for Simple Question Answering over Knowledge Graphs
- An Evaluation of Transfer Learning for Classifying Sales Engagement Emails at Large Scale