Using Fast Weights to Attend to the Recent Past
arXiv:1610.06258
Abstract
Until recently, research on artificial neural networks was largely restricted to systems with only two types of variable: Neural activities that represent the current or recent input and weights that learn to capture regularities among inputs, outputs and payoffs. There is no good reason for this restriction. Synapses have dynamics at many different time-scales and this suggests that artificial neural networks might benefit from variables that change slower than activities but much faster than the standard weights. These "fast weights" can be used to store temporary memories of the recent past and they provide a neurally plausible way of implementing the type of attention to the past that has recently proved very helpful in sequence-to-sequence models. By using fast weights we can avoid the need to store copies of neural activity patterns.
Added [Schmidhuber 1993] citation to the last paragraph of the introduction. Fixed typo appendix A.1 uniform initialization to 1/\sqrt{H}
Cited by in corpus (25)
- Artificial neural networks for neuroscientists: A primer
- Random Feature Attention
- Large-Scale Long-Tailed Recognition in an Open World
- On the Binding Problem in Artificial Neural Networks
- Evolving Modular Soft Robots without Explicit Inter-Module Communication using Local Self-Attention
- Emergent Symbols through Binding in External Memory
- Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
- The Sensory Neuron as a Transformer: Permutation-Invariant Neural Networks for Reinforcement Learning
- Learning to update Auto-associative Memory in Recurrent Neural Networks for Improving Sequence Memorization
- Learning to Learn with Feedback and Local Plasticity
- Dynamic Memory Induction Networks for Few-Shot Text Classification
- Few-shot Sequence Learning with Transformers
- Continual Learning with Deep Artificial Neurons
- A Biologically Inspired Visual Working Memory for Deep Networks
- Contextual Recurrent Neural Networks
- Variational Tracking and Prediction with Generative Disentangled State-Space Models
- Modeling Past and Future for Neural Machine Translation
- Rotational Unit of Memory
- Learning Associative Inference Using Fast Weight Memory
- A Field Guide to Scientific XAI: Transparent and Interpretable Deep Learning for Bioinformatics Research
- A Memory-Augmented Neural Network Model of Abstract Rule Learning
- Memory and attention in deep learning
- Representation Memorization for Fast Learning New Knowledge without Forgetting
- Back to Square One: Superhuman Performance in Chutes and Ladders Through Deep Neural Networks and Tree Search
- Beyond Short Clips: End-to-End Video-Level Learning with Collaborative Memories