A Joint Many-Task Model: Growing a Neural Network for Multiple NLP Tasks
arXiv:1611.01587
Abstract
Transfer and multi-task learning have traditionally focused on either a single source-target pair or very few, similar tasks. Ideally, the linguistic levels of morphology, syntax and semantics would benefit each other by being trained in a single model. We introduce a joint many-task model together with a strategy for successively growing its depth to solve increasingly complex tasks. Higher layers include shortcut connections to lower-level task predictions to reflect linguistic hierarchies. We use a simple regularization term to allow for optimizing all model weights to improve one task's loss without exhibiting catastrophic interference of the other tasks. Our single end-to-end model obtains state-of-the-art or competitive results on five different tasks from tagging, parsing, relatedness, and entailment tasks.
Accepted as a full paper at the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP 2017)
References in corpus (7)
- Sequence to Sequence Learning with Neural Networks
- Transition-Based Dependency Parsing with Stack Long Short-Term Memory
- Gated Word-Character Recurrent Language Model
- Charagram: Embedding Words and Sentences via Character n-grams
- Tree-to-Sequence Attentional Neural Machine Translation
- Structured Training for Neural Network Transition-Based Parsing
- Modelling Sentence Pairs with Tree-structured Attentive Encoder
Cited by in corpus (30)
- An Overview of Multi-Task Learning in Deep Neural Networks
- CTRL: A Conditional Transformer Language Model for Controllable Generation
- GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks
- Multi-Task Learning with Deep Neural Networks: A Survey
- The Natural Language Decathlon: Multitask Learning as Question Answering
- Optimal Hyperparameters for Deep LSTM-Networks for Sequence Labeling Tasks
- Glyce: Glyph-vectors for Chinese Character Representations
- Learning General Purpose Distributed Sentence Representations via Large Scale Multi-task Learning
- Transferable Multi-Domain State Generator for Task-Oriented Dialogue Systems
- Shortcut-Stacked Sentence Encoders for Multi-Domain Inference
- Multi-task Deep Reinforcement Learning with PopArt
- Beyond Shared Hierarchies: Deep Multitask Learning through Soft Layer Ordering
- Getting To Know You: User Attribute Extraction from Dialogues
- Semantic Parsing with Syntax- and Table-Aware SQL Generation
- Improving Interpretability of Word Embeddings by Generating Definition and Usage
- Evolutionary Architecture Search For Deep Multitask Networks
- Multi-Label Transfer Learning for Multi-Relational Semantic Similarity
- Exploring the Syntactic Abilities of RNNs with Multi-task Learning
- HyperGrid: Efficient Multi-Task Transformers with Grid-wise Decomposable Hyper Projections
- TextNAS: A Neural Architecture Search Space tailored for Text Representation
- Multi-task Neural Network for Non-discrete Attribute Prediction in Knowledge Graphs
- Linguistically-Enriched and Context-Aware Zero-shot Slot Filling
- Empirical Evaluation of Multi-task Learning in Deep Neural Networks for Natural Language Processing
- Graph-Driven Generative Models for Heterogeneous Multi-Task Learning
- Hierarchical Multitask Learning Approach for BERT
- Copy-Enhanced Heterogeneous Information Learning for Dialogue State Tracking
- A Hierarchical Deep Learning Natural Language Parser for Fashion
- SAFE: Spectral Evolution Analysis Feature Extraction for Non-Stationary Time Series Prediction
- Using Context Information to Enhance Simple Question Answering
- Improving Limited Labeled Dialogue State Tracking with Self-Supervision