Attention Approximates Sparse Distributed Memory
arXiv:2111.05498
Abstract
While Attention has come to be an important mechanism in deep learning, there remains limited intuition for why it works so well. Here, we show that Transformer Attention can be closely related under certain data conditions to Kanerva's Sparse Distributed Memory (SDM), a biologically plausible associative memory model. We confirm that these conditions are satisfied in pre-trained GPT2 Transformer models. We discuss the implications of the Attention-SDM map and provide new computational and biological interpretations of Attention.
References in corpus (10)
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- Language Models are Few-Shot Learners
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
- Transformers are RNNs: Fast Autoregressive Transformers with Linear Attention
- Rethinking Attention with Performers
- Visualizing Attention in Transformer-Based Language Representation Models
- Linear Transformers Are Secretly Fast Weight Programmers
- Are Convolutional Neural Networks or Transformers more like human vision?
- Do Transformer Attention Heads Provide Transparency in Abstractive Summarization?
- Product Kanerva Machines: Factorized Bayesian Memory