1 citations · 2 across the 9 of their papers we have counts for
9 papers · 1 filter
Universal Transformers for Circuit Computations: Perfect Length Generalization in Tiny Transformers
Takuya Ito, Ruchir Puri, Murray Campbell +1
Learning generalizable algorithmic computations remains a challenge for neural networks, as reflected in persistent failures on compositional and length generalization benchmarks.…
What Time Is It? How Data Geometry Makes Time Conditioning Optional for Flow Matching
Alec Helbling, Sebastian Gutierrez Hernandez, Benjamin Hoover +2
Recent work has shown that models flow matching models can be trained without explicit time conditioning, challenging the standard view that the interpolation time is needed to dis…
Modern Methods in Associative Memory
Dmitry Krotov, Benjamin Hoover, Parikshit Ram +1
Associative Memories like the famous Hopfield Networks are elegant models for describing fully recurrent neural networks whose fundamental job is to store and retrieve information.…
Transformers Learn Faster with Semantic Focus
Parikshit Ram, Kenneth L. Clarkson, Tim Klinger +2
Various forms of sparse attention have been explored to mitigate the quadratic computational and memory cost of the attention mechanism in transformers. We study sparse transformer…
Transformer Circuits Can Realize Clustering Algorithms
Kenneth L. Clarkson, Lior Horesh, Takuya Ito +2
Although transformers are most commonly optimized as statistical sequence models, it is unclear to what extent they can implement and learn exact algorithmic computations. Here, we…
Dense Associative Memory with Epanechnikov Energy
Benjamin Hoover, Zhaoyang Shi, Krishnakumar Balasubramanian +2
We propose a novel energy function for Dense Associative Memory (DenseAM) networks, the log-sum-ReLU (LSR), inspired by optimal kernel density estimation. Unlike the common log-sum…