13 citations · 15 across the 2 of their papers we have counts for
5 papers
Going Beyond Linear Transformers with Recurrent Fast Weight Programmers
Kazuki Irie, Imanol Schlag, Róbert Csordás +1
Transformers with linearised attention (''linear Transformers'') have demonstrated the practical scalability and effectiveness of outer product-based Fast Weight Programmers (FWPs)…
Linear Transformers Are Secretly Fast Weight Programmers
Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber
We show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early '90s, where a ``slow" neural net learns by gradient descent to pro…
Learning Associative Inference Using Fast Weight Memory
Imanol Schlag, Tsendsuren Munkhdalai, Jürgen Schmidhuber
Humans can quickly associate stimuli to solve problems in novel contexts. Our novel neural network model learns state representations of facts that can be composed to perform such…
Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving
Imanol Schlag, Paul Smolensky, Roland Fernandez +3
We incorporate Tensor-Product Representations within the Transformer in order to better support the explicit representation of relation structure. Our Tensor-Product Transformer (T…
Learning to Reason with Third-Order Tensor Products
Imanol Schlag, Jürgen Schmidhuber
We combine Recurrent Neural Networks with Tensor Product Representations to learn combinatorial representations of sequential data. This improves symbolic interpretation and system…