Frosting Weights for Better Continual Training
arXiv:2001.01829 · doi:10.1109/ICMLA.2019.00094
Abstract
Training a neural network model can be a lifelong learning process and is a computationally intensive one. A severe adverse effect that may occur in deep neural network models is that they can suffer from catastrophic forgetting during retraining on new data. To avoid such disruptions in the continuous learning, one appealing property is the additive nature of ensemble models. In this paper, we propose two generic ensemble approaches, gradient boosting and meta-learning, to solve the catastrophic forgetting problem in tuning pre-trained neural network models.
References in corpus (11)
- Distilling the Knowledge in a Neural Network
- How transferable are features in deep neural networks?
- An Overview of Multi-Task Learning in Deep Neural Networks
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- Efficient Lifelong Learning with A-GEM
- On Tiny Episodic Memories in Continual Learning
- Slimmable Neural Networks
- Continual Learning in Generative Adversarial Nets
- Improving and Understanding Variational Continual Learning
- Model Distillation with Knowledge Transfer from Face Classification to Alignment and Verification
- Benchmark Environments for Multitask Learning in Continuous Domains