4 papers
Sparse Attention as Compact Kernel Regression
Saul Santos, Nuno Gonçalves, Daniel C. McNamee +2
Recent work has revealed a link between self-attention mechanisms in transformers and test-time kernel regression via the Nadaraya-Watson estimator, with standard softmax attention…
Hopfield-Fenchel-Young Networks: A Unified Framework for Associative Memory Retrieval
Saul Santos, Vlad Niculae, Daniel McNamee +1
Associative memory models, such as Hopfield networks and their modern variants, have garnered renewed interest due to advancements in memory capacity and connections with self-atte…
-Video: A Training-Free Approach to Long Video Understanding via Continuous-Time Memory Consolidation
Saul Santos, António Farinhas, Daniel C. McNamee +1
Current video-language models struggle with long-video understanding due to limited context lengths and reliance on sparse frame subsampling, often leading to information loss. Thi…
Sparse and Structured Hopfield Networks
Saul Santos, Vlad Niculae, Daniel McNamee +1
Modern Hopfield networks have enjoyed recent interest due to their connection to attention in transformers. Our paper provides a unified framework for sparse Hopfield networks by e…