Overcoming catastrophic forgetting with hard attention to the task
arXiv:1801.01423
Abstract
Catastrophic forgetting occurs when a neural network loses the information learned in a previous task after training on subsequent tasks. This problem remains a hurdle for artificial intelligence systems with sequential learning capabilities. In this paper, we propose a task-based hard attention mechanism that preserves previous tasks' information without affecting the current task's learning. A hard attention mask is learned concurrently to every task, through stochastic gradient descent, and previous masks are exploited to condition such learning. We show that the proposed mechanism is effective for reducing catastrophic forgetting, cutting current rates by 45 to 80%. We also show that it is robust to different hyperparameter choices, and that it offers a number of monitoring capabilities. The approach features the possibility to control both the stability and compactness of the learned knowledge, which we believe makes it also attractive for online learning or network compression applications.
Includes appendix. Accepted for ICML 2018
References in corpus (7)
- Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
- PathNet: Evolution Channels Gradient Descent in Super Neural Networks
- Variational Continual Learning
- Less-forgetting Learning in Deep Neural Networks
- On the Origin of Deep Learning
- Memory-based Parameter Adaptation
- Overcoming Catastrophic Interference by Conceptors
Cited by in corpus (53)
- A continual learning survey: Defying forgetting in classification tasks
- Three scenarios for continual learning
- Alleviating catastrophic forgetting using context-dependent gating and synaptic stabilization
- Generative replay with feedback connections as a general strategy for continual learning
- Incremental Object Detection via Meta-Learning
- Learning to Learn without Forgetting by Maximizing Transfer and Minimizing Interference
- Don't forget, there is more than forgetting: new metrics for Continual Learning
- Learn to Grow: A Continual Structure Learning Framework for Overcoming Catastrophic Forgetting
- Selfless Sequential Learning
- Triple Memory Networks: a Brain-Inspired Method for Continual Learning
- Superposition of many models into one
- Rotate your Networks: Better Weight Consolidation and Less Catastrophic Forgetting
- An Investigation of Replay-based Approaches for Continual Learning
- Efficient Continual Learning with Modular Networks and Task-Driven Priors
- SpaRCe: Improved Learning of Reservoir Computing Systems through Sparse Representations
- Class-incremental Learning via Deep Model Consolidation
- Meta-Consolidation for Continual Learning
- Continual Learning with Node-Importance based Adaptive Group Sparse Regularization
- Using Adapters to Overcome Catastrophic Forgetting in End-to-End Automatic Speech Recognition
- Multi-View Class Incremental Learning
- Few-Shot Class-Incremental Learning
- A Neural Span-Based Continual Named Entity Recognition Model
- Generalisation Guarantees for Continual Learning with Orthogonal Gradient Descent
- Continual Learning with Knowledge Transfer for Sentiment Classification
- Autoencoder-Based Incremental Class Learning without Retraining on Old Data
- Self-Supervised Learning Aided Class-Incremental Lifelong Learning
- Distribution Aligned Semantics Adaption for Lifelong Person Re-Identification
- Energy-Based Models for Continual Learning
- Lifelong Neural Predictive Coding: Learning Cumulatively Online without Forgetting
- CUCL: Codebook for Unsupervised Continual Learning
- Continual Learning Using Multi-view Task Conditional Neural Networks
- Overcoming Catastrophic Forgetting by Neuron-level Plasticity Control
- Sequoia: A Software Framework to Unify Continual Learning Research
- Continual Learning for Text Classification with Information Disentanglement Based Regularization
- Bayesian Structure Adaptation for Continual Learning
- Single-Net Continual Learning with Progressive Segmented Training (PST)
- Continual Learning via Bit-Level Information Preserving
- Frosting Weights for Better Continual Training
- Local learning rules to attenuate forgetting in neural networks
- Forget Me Not: Reducing Catastrophic Forgetting for Domain Adaptation in Reading Comprehension
- Balancing Specialization, Generalization, and Compression for Detection and Tracking
- Continual Learning: Forget-free Winning Subnetworks for Video Representations
- Bilevel Continual Learning
- Dynamic Continual Learning: Harnessing Parameter Uncertainty for Improved Network Adaptation
- Continual Learning With Quasi-Newton Methods
- Group and Exclusive Sparse Regularization-based Continual Learning of CNNs
- TAG: Task-based Accumulated Gradients for Lifelong learning
- Split-and-Bridge: Adaptable Class Incremental Learning within a Single Neural Network
- Class Incremental Online Streaming Learning
- Online Continual Learning in Image Classification: An Empirical Survey
- A Deep Optimization Approach for Image Deconvolution
- Continual learning using hash-routed convolutional neural networks
- LIRA: Lifelong Image Restoration from Unknown Blended Distortions