Pseudo-task Augmentation: From Deep Multitask Learning to Intratask Sharing---and Back
arXiv:1803.04062
Abstract
Deep multitask learning boosts performance by sharing learned structure across related tasks. This paper adapts ideas from deep multitask learning to the setting where only a single task is available. The method is formalized as pseudo-task augmentation, in which models are trained with multiple decoders for each task. Pseudo-tasks simulate the effect of training towards closely-related tasks drawn from the same universe. In a suite of experiments, pseudo-task augmentation is shown to improve performance on single-task learning problems. When combined with multitask learning, further improvements are achieved, including state-of-the-art performance on the CelebA dataset, showing that pseudo-task augmentation and multitask learning have complementary value. All in all, pseudo-task augmentation is a broadly applicable and efficient way to boost performance in deep learning systems.
Published as a conference paper at ICML 2018; 10 pages
References in corpus (9)
- Practical Bayesian Optimization of Machine Learning Algorithms
- Neural Architecture Search with Reinforcement Learning
- An Overview of Multi-Task Learning in Deep Neural Networks
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- Learning multiple visual domains with residual adapters
- Gradient-based Hyperparameter Optimization through Reversible Learning
- One Model To Learn Them All
- Learned in Translation: Contextualized Word Vectors
- Distral: Robust Multitask Reinforcement Learning
Cited by in corpus (8)
- Neural Architecture Search: A Survey
- Just Pick a Sign: Optimizing Deep Multitask Models with Gradient Sign Dropout
- Deep Multi-Task Learning for Malware Image Classification
- Optimizing Airbnb Search Journey with Multi-task Learning
- Small Towers Make Big Differences
- Modular Universal Reparameterization: Deep Multi-task Learning Across Diverse Domains
- Frosting Weights for Better Continual Training
- Exploring Correlations in Multiple Facial Attributes through Graph Attention Network