An Empirical Investigation of Catastrophic Forgetting in Gradient-Based Neural Networks
arXiv:1312.6211
Abstract
Catastrophic forgetting is a problem faced by many machine learning models and algorithms. When trained on one task, then trained on a second task, many machine learning models "forget" how to perform the first task. This is widely believed to be a serious problem for neural networks. Here, we investigate the extent to which the catastrophic forgetting problem occurs for modern neural networks, comparing both established and recent gradient-based training algorithms and activation functions. We also examine the effect of the relationship between the first task and the second task on catastrophic forgetting. We find that it is always best to train using the dropout algorithm--the dropout algorithm is consistently best at adapting to the new task, remembering the old task, and has the best tradeoff curve between these two extremes. We find that different tasks and relationships between tasks result in very different rankings of activation function performance. This suggests the choice of activation function should always be cross-validated.
References in corpus (1)
Cited by in corpus (205)
- Overcoming catastrophic forgetting in neural networks
- A continual learning survey: Defying forgetting in classification tasks
- How To Backdoor Federated Learning
- A Survey and Critique of Multiagent Deep Reinforcement Learning
- Three scenarios for continual learning
- A Closer Look at Memorization in Deep Networks
- Three Approaches for Personalization with Applications to Federated Learning
- FearNet: Brain-Inspired Model for Incremental Learning
- Alleviating catastrophic forgetting using context-dependent gating and synaptic stabilization
- Off-Policy Deep Reinforcement Learning without Exploration
- Overcoming catastrophic forgetting with hard attention to the task
- Re-evaluating Continual Learning Scenarios: A Categorization and Case for Strong Baselines
- Continual Learning with Deep Generative Replay
- Generative replay with feedback connections as a general strategy for continual learning
- Towards Robust Evaluations of Continual Learning
- Less-forgetting Learning in Deep Neural Networks
- Incremental Learning for Semantic Segmentation of Large-Scale Remote Sensing Data
- SelfHAR: Improving Human Activity Recognition through Self-training with Unlabeled Data
- Deep Generative Dual Memory Network for Continual Learning
- Variational Federated Multi-Task Learning
- Coresets via Bilevel Optimization for Continual Learning and Streaming
- Functional Regularisation for Continual Learning with Gaussian Processes
- Online Continual Learning with Maximally Interfered Retrieval
- Selfless Sequential Learning
- ASP: Learning to Forget with Adaptive Synaptic Plasticity in Spiking Neural Networks
- Task Agnostic Continual Learning via Meta Learning
- Cognitively-Inspired Model for Incremental Learning Using a Few Examples
- Superposition of many models into one
- Lifelong Machine Learning Potentials
- CLeaR: An Adaptive Continual Learning Framework for Regression Tasks
- Subspace Regularizers for Few-Shot Class Incremental Learning
- M2KD: Multi-model and Multi-level Knowledge Distillation for Incremental Learning
- A Strategy for an Uncompromising Incremental Learner
- DeepSpace: An Online Deep Learning Framework for Mobile Big Data to Understand Human Mobility Patterns
- Diffusion-based neuromodulation can eliminate catastrophic forgetting in simple neural networks
- Orthogonal Gradient Descent for Continual Learning
- End-to-End Incremental Learning
- Class-incremental Learning with Pre-allocated Fixed Classifiers
- Toward Understanding Catastrophic Forgetting in Continual Learning
- Federated Learning with Additional Mechanisms on Clients to Reduce Communication Costs
- Improving and Understanding Variational Continual Learning
- Continual Learning in Low-rank Orthogonal Subspaces
- Low-shot Learning via Covariance-Preserving Adversarial Augmentation Networks
- Towards Making the Most of BERT in Neural Machine Translation
- Efficient Continual Learning with Modular Networks and Task-Driven Priors
- Controllable Semantic Parsing via Retrieval Augmentation
- FedAT: A High-Performance and Communication-Efficient Federated Learning System with Asynchronous Tiers
- Class-incremental Learning via Deep Model Consolidation
- Defining Benchmarks for Continual Few-Shot Learning
- Laplace Redux -- Effortless Bayesian Deep Learning
- Continual Learning with Adaptive Weights (CLAW)
- Toward Continual Learning for Conversational Agents
- MetricGAN+: An Improved Version of MetricGAN for Speech Enhancement
- RATT: Recurrent Attention to Transient Tasks for Continual Image Captioning
- Meta Continual Learning
- Lifelong Object Detection
- Continual Learning via Inter-Task Synaptic Mapping
- Active Long Term Memory Networks
- Anatomy of Catastrophic Forgetting: Hidden Representations and Task Semantics
- Continual Reinforcement Learning with Complex Synapses
- Catastrophic forgetting: still a problem for DNNs
- Preserving Earlier Knowledge in Continual Learning with the Help of All Previous Feature Extractors
- Multi-Attribute Selectivity Estimation Using Deep Learning
- Rainbow Memory: Continual Learning with a Memory of Diverse Samples
- Interoceptive robustness through environment-mediated morphological development
- Continual Classification Learning Using Generative Models
- Two-Level Residual Distillation based Triple Network for Incremental Object Detection
- Unsupervised Pretraining for Sequence to Sequence Learning
- Training Binary Neural Networks using the Bayesian Learning Rule
- Quantum Continual Learning Overcoming Catastrophic Forgetting
- Experience Replay Using Transition Sequences
- Physics-Guided Continual Learning for Predicting Emerging Aqueous Organic Redox Flow Battery Material Performance
- A Theoretical Analysis of Catastrophic Forgetting through the NTK Overlap Matrix
- Using Small Proxy Datasets to Accelerate Hyperparameter Search
- CLAR: Contrastive Learning of Auditory Representations
- GP-Tree: A Gaussian Process Classifier for Few-Shot Incremental Learning
- Transferring Autonomous Driving Knowledge on Simulated and Real Intersections
- Continual Learning in Neural Networks
- Bootstrap Representation Learning for Segmentation on Medical Volumes and Sequences
- Online Fast Adaptation and Knowledge Accumulation: a New Approach to Continual Learning
- Deep learning via message passing algorithms based on belief propagation
- Linear Mode Connectivity in Multitask and Continual Learning
- Advances in Continual Graph Learning for Anti-Money Laundering Systems: A Comprehensive Review
- Behavior From the Void: Unsupervised Active Pre-Training
- Generalisation Guarantees for Continual Learning with Orthogonal Gradient Descent
- Analyzing Knowledge Transfer in Deep Q-Networks for Autonomously Handling Multiple Intersections
- Keep and Learn: Continual Learning by Constraining the Latent Space for Knowledge Preservation in Neural Networks
- Recall and Learn: Fine-tuning Deep Pretrained Language Models with Less Forgetting
- An Incremental Learning framework for Large-scale CTR Prediction
- Self-Supervised Learning to Prove Equivalence Between Straight-Line Programs via Rewrite Rules
- Encoders and Ensembles for Task-Free Continual Learning
- Hierarchical Indian Buffet Neural Networks for Bayesian Continual Learning
- ContCap: A scalable framework for continual image captioning
- Domain Adaptive Text Style Transfer
- Organizing Experience: A Deeper Look at Replay Mechanisms for Sample-based Planning in Continuous State Domains
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image Retrieval
- Unpacking Information Bottlenecks: Unifying Information-Theoretic Objectives in Deep Learning
- Recurrent Spectral Network (RSN): shaping the basin of attraction of a discrete map to reach automated classification
- Never Forget: Balancing Exploration and Exploitation via Learning Optical Flow
- A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition
- Supermasks in Superposition
- Adaptive Online Planning for Continual Lifelong Learning
- Continual Learning using a Bayesian Nonparametric Dictionary of Weight Factors
- Improving Performance in Reinforcement Learning by Breaking Generalization in Neural Networks
- Domain Adaptive Medical Image Segmentation via Adversarial Learning of Disease-Specific Spatial Patterns
- Dissecting Catastrophic Forgetting in Continual Learning by Deep Visualization
- Disentangled Representations in Neural Models
- Boosting Binary Masks for Multi-Domain Learning through Affine Transformations
- Meta-learnt priors slow down catastrophic forgetting in neural networks
- Language Models for German Text Simplification: Overcoming Parallel Data Scarcity through Style-specific Pre-training
- Artificial Neural Variability for Deep Learning: On Overfitting, Noise Memorization, and Catastrophic Forgetting
- Robust Importance Sampling for Error Estimation in the Context of Optimal Bayesian Transfer Learning
- Evaluating Online Continual Learning with CALM
- Molecular De Novo Design through Deep Reinforcement Learning
- Powerpropagation: A sparsity inducing weight reparameterisation
- Regularizing Towards Permutation Invariance in Recurrent Models
- Continuous Coordination As a Realistic Scenario for Lifelong Learning
- Faster ILOD: Incremental Learning for Object Detectors based on Faster RCNN
- On the Importance of Difficulty Calibration in Membership Inference Attacks
- Concept-Oriented Deep Learning
- Conditional Computation for Continual Learning
- Learning Vision-based Robotic Manipulation Tasks Sequentially in Offline Reinforcement Learning Settings
- On the importance of cross-task features for class-incremental learning
- Continual Representation Learning for Biometric Identification
- Continual Learning in Deep Neural Network by Using a Kalman Optimiser
- On the Costs and Benefits of Adopting Lifelong Learning for Software Analytics -- Empirical Study on Brown Build and Risk Prediction
- Localizing Catastrophic Forgetting in Neural Networks
- Positive-Congruent Training: Towards Regression-Free Model Updates
- OpenLORIS-Object: A Robotic Vision Dataset and Benchmark for Lifelong Deep Learning
- Continual Learning via Online Leverage Score Sampling
- Unifying Regularisation Methods for Continual Learning
- A Comparative Study of Calibration Methods for Imbalanced Class Incremental Learning
- A Saak Transform Approach to Efficient, Scalable and Robust Handwritten Digits Recognition
- Continual Learning Using Bayesian Neural Networks
- Wide Neural Networks Forget Less Catastrophically
- Posterior Meta-Replay for Continual Learning
- A Combinatorial Perspective on Transfer Learning
- Learning a Multi-Domain Curriculum for Neural Machine Translation
- Continual Learning Using Multi-view Task Conditional Neural Networks
- Unsupervised Transfer Learning with Self-Supervised Remedy
- Deep Learning for RF Signal Classification in Unknown and Dynamic Spectrum Environments
- ExpertMatcher: Automating ML Model Selection for Users in Resource Constrained Countries
- Effective Data Augmentation with Multi-Domain Learning GANs
- Continual Learning with Self-Organizing Maps
- End-to-End 3D-PointCloud Semantic Segmentation for Autonomous Driving
- Incremental Learning for Personalized Recommender Systems
- Less-forgetful Learning for Domain Expansion in Deep Neural Networks
- Continual Learning: Tackling Catastrophic Forgetting in Deep Neural Networks with Replay Processes
- A Foliated View of Transfer Learning
- Rethinking Experience Replay: a Bag of Tricks for Continual Learning
- Overcoming Catastrophic Forgetting by Soft Parameter Pruning
- Learn-Prune-Share for Lifelong Learning
- Unsupervised Pre-training for Natural Language Generation: A Literature Review
- Review, Analysis and Design of a Comprehensive Deep Reinforcement Learning Framework
- Knowledge transfer in deep block-modular neural networks
- Memory Efficient Class-Incremental Learning for Image Classification
- FDCNet: Feature Drift Compensation Network for Class-Incremental Weakly Supervised Object Localization
- Learn Faster and Forget Slower via Fast and Stable Task Adaptation
- The unreasonable effectiveness of Batch-Norm statistics in addressing catastrophic forgetting across medical institutions
- Neural Machine Translation: A Review and Survey
- Problems in AI research and how the SP System may help to solve them
- Graph-Based Continual Learning
- Adaptive Compression-based Lifelong Learning
- Learning to Remember from a Multi-Task Teacher
- Power Law in Sparsified Deep Neural Networks
- Explaining How Deep Neural Networks Forget by Deep Visualization
- Self-Net: Lifelong Learning via Continual Self-Modeling
- Towards Recognizing New Semantic Concepts in New Visual Domains
- Reducing Catastrophic Forgetting in Modular Neural Networks by Dynamic Information Balancing
- Differentially Private Distributed Learning for Language Modeling Tasks
- Collaborative Semantic Aggregation and Calibration for Federated Domain Generalization
- Revisiting Catastrophic Forgetting in Class Incremental Learning
- The LMU Munich System for the WMT 2020 Unsupervised Machine Translation Shared Task
- Compositional pre-training for neural semantic parsing
- Questions to Guide the Future of Artificial Intelligence Research
- Learning to Transfer: A Foliated Theory
- Statistical Mechanical Analysis of Catastrophic Forgetting in Continual Learning with Teacher and Student Networks
- A good body is all you need: avoiding catastrophic interference via agent architecture search
- Weight Friction: A Simple Method to Overcome Catastrophic Forgetting and Enable Continual Learning
- Wav-BERT: Cooperative Acoustic and Linguistic Representation Learning for Low-Resource Speech Recognition
- BooVAE: Boosting Approach for Continual Learning of VAE
- Online Continual Learning in Image Classification: An Empirical Survey
- Iterative Semi-parametric Dynamics Model Learning For Autonomous Racing
- Lethean Attack: An Online Data Poisoning Technique
- Federated Marginal Personalization for ASR Rescoring
- A Biologically Plausible Audio-Visual Integration Model for Continual Learning
- CALM: Continuous Adaptive Learning for Language Modeling
- Does the Adam Optimizer Exacerbate Catastrophic Forgetting?
- A Review of Computer Vision Methods in Network Security
- Rethinking Neural Networks With Benford's Law
- Applying Incremental Deep Neural Networks-based Posture Recognition Model for Injury Risk Assessment in Construction
- iLGaCo: Incremental Learning of Gait Covariate Factors
- Using Fictitious Class Representations to Boost Discriminative Zero-Shot Learners
- Behavioral Experiments for Understanding Catastrophic Forgetting
- TAG: Task-based Accumulated Gradients for Lifelong learning
- Batch Normalization and the impact of batch structure on the behavior of deep convolution networks
- Learning Multiple Categories on Deep Convolution Networks
- Large-scale Kernel Methods and Applications to Lifelong Robot Learning
- SeNA-CNN: Overcoming Catastrophic Forgetting in Convolutional Neural Networks by Selective Network Augmentation
- Incremental Class Learning using Variational Autoencoders with Similarity Learning
- Contrast R-CNN for Continual Learning in Object Detection
- Solving hybrid machine learning tasks by traversing weight space geodesics
- Continual One-Shot Learning of Hidden Spike-Patterns with Neural Network Simulation Expansion and STDP Convergence Predictions
- Data augmentation and pre-trained networks for extremely low data regimes unsupervised visual inspection
- Embodiment dictates learnability in neural controllers