activity
20182021
most citedGoing Beyond Linear Transformers with Recurrent Fast Weight Programmers

13 citations · 15 across the 2 of their papers we have counts for

collaborators

5 papers

cs.LG202113 cited

Going Beyond Linear Transformers with Recurrent Fast Weight Programmers

Kazuki Irie, Imanol Schlag, Róbert Csordás +1

Transformers with linearised attention (''linear Transformers'') have demonstrated the practical scalability and effectiveness of outer product-based Fast Weight Programmers (FWPs)…

cs.LG2021

Linear Transformers Are Secretly Fast Weight Programmers

Imanol Schlag, Kazuki Irie, Jürgen Schmidhuber

We show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early '90s, where a ``slow" neural net learns by gradient descent to pro…

cs.LG20202 cited

Learning Associative Inference Using Fast Weight Memory

Imanol Schlag, Tsendsuren Munkhdalai, Jürgen Schmidhuber

Humans can quickly associate stimuli to solve problems in novel contexts. Our novel neural network model learns state representations of facts that can be composed to perform such…

cs.LG2019

Enhancing the Transformer with Explicit Relational Encoding for Math Problem Solving

Imanol Schlag, Paul Smolensky, Roland Fernandez +3

We incorporate Tensor-Product Representations within the Transformer in order to better support the explicit representation of relation structure. Our Tensor-Product Transformer (T…

cs.LG2018

Learning to Reason with Third-Order Tensor Products

Imanol Schlag, Jürgen Schmidhuber

We combine Recurrent Neural Networks with Tensor Product Representations to learn combinatorial representations of sequential data. This improves symbolic interpretation and system…