Hopfield Networks is All You Need
arXiv:2008.02217
Abstract
We introduce a modern Hopfield network with continuous states and a corresponding update rule. The new Hopfield network can store exponentially (with the dimension of the associative space) many patterns, retrieves the pattern with one update, and has exponentially small retrieval errors. It has three types of energy minima (fixed points of the update): (1) global fixed point averaging over all patterns, (2) metastable states averaging over a subset of patterns, and (3) fixed points which store a single pattern. The new update rule is equivalent to the attention mechanism used in transformers. This equivalence enables a characterization of the heads of transformer models. These heads perform in the first layers preferably global averaging and in higher layers partial averaging via metastable states. The new modern Hopfield network can be integrated into deep learning architectures as layers to allow the storage of and access to raw input data, intermediate results, or learned prototypes. These Hopfield layers enable new ways of deep learning, beyond fully-connected, convolutional, or recurrent networks, and provide pooling, memory, association, and attention mechanisms. We demonstrate the broad applicability of the Hopfield layers across various domains. Hopfield layers improved state-of-the-art on three out of four considered multiple instance learning problems as well as on immune repertoire classification with several hundreds of thousands of instances. On the UCI benchmark collections of small classification tasks, where deep learning methods typically struggle, Hopfield layers yielded a new state-of-the-art when compared to different machine learning methods. Finally, Hopfield layers achieved state-of-the-art on two drug design datasets. The implementation is available at: https://github.com/ml-jku/hopfield-layers
10 pages (+ appendix); 12 figures; Blog: https://ml-jku.github.io/hopfield-layers/; GitHub: https://github.com/ml-jku/hopfield-layers
References in corpus (14)
- Deep Learning in Neural Networks: An Overview
- PyTorch: An Imperative Style, High-Performance Deep Learning Library
- Semi-Supervised Classification with Graph Convolutional Networks
- HuggingFace's Transformers: State-of-the-art Natural Language Processing
- Pointer Sentinel Mixture Models
- Synthesizer: Rethinking Self-Attention in Transformer Models
- Using Fast Weights to Attend to the Recent Past
- Topological and Dynamical Complexity of Random Neural Networks
- Large Associative Memory Problem in Neurobiology and Machine Learning
- Permutation-equivariant neural networks applied to dynamics prediction
- Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving
- Learning to update Auto-associative Memory in Recurrent Neural Networks for Improving Sequence Memorization
- Encoding-based Memory Modules for Recurrent Neural Networks
- Set Distribution Networks: a Generative Model for Sets of Images
Cited by in corpus (31)
- Pretrained Transformers as Universal Computation Engines
- Language Models are Open Knowledge Graphs
- Large Associative Memory Problem in Neurobiology and Machine Learning
- Hypercomplex-Valued Recurrent Correlation Neural Networks
- Cross-Domain Few-Shot Learning by Representation Fusion
- Centroid Transformers: Learning to Abstract with Attention
- RealFormer: Transformer Likes Residual Attention
- EarlyBERT: Efficient BERT Training via Early-bird Lottery Tickets
- Hierarchical Associative Memory
- Eight challenges in developing theory of intelligence
- Scatterbrain: Unifying Sparse and Low-rank Attention Approximation
- A deep learning theory for neural networks grounded in physics
- Dense Hopfield Networks in the Teacher-Student Setting
- Trusted Artificial Intelligence: Towards Certification of Machine Learning Applications
- Supervising the Transfer of Reasoning Patterns in VQA
- COVID-19 Pneumonia Severity Prediction using Hybrid Convolution-Attention Neural Architectures
- A generalized Hopfield model to store and retrieve mismatched memory patterns
- Unitary Evolutions Sourced By Interacting Quantum Memories: Closed Quantum Systems Directing Themselves Using Their State Histories
- Modern Hopfield Networks for Few- and Zero-Shot Reaction Template Prediction
- A self-learning magnetic Hopfield neural network with intrinsic gradient descent adaption
- On the Distribution, Sparsity, and Inference-time Quantization of Attention Values in Transformers
- Restricted Boltzmann Machine and Deep Belief Network: Tutorial and Survey
- Learning Associative Inference Using Fast Weight Memory
- Overcoming Quadratic Hardware Scaling for a Fully Connected Digital Oscillatory Neural Network
- HiCOMEX: Facial Action Unit Recognition Based on Hierarchy Intensity Distribution and COMEX Relation Learning
- Probabilistic Transformers
- Simplified derivations for high-dimensional convex learning problems
- Critical Dynamics and Cyclic Memory Retrieval in Non-reciprocal Hopfield Networks
- Focus on the present: a regularization method for the ASR source-target attention layer
- The Koha Code: A Biological Theory of Memory
- HBert + BiasCorp -- Fighting Racism on the Web