45 citations · 74 across the 6 of their papers we have counts for
6 papers · 1 filter
Learning State-Tracking from Code Using Linear RNNs
Julien Siems, Riccardo Grazzi, Korbinian Pöppel +3
Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence models architectures like Transformers a…
DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products
Julien Siems, Timur Carstensen, Arber Zela +3
Linear Recurrent Neural Networks (linear RNNs) have emerged as competitive alternatives to Transformers for sequence modeling, offering efficient training and linear-time inference…
Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues
Riccardo Grazzi, Julien Siems, Arber Zela +3
Linear Recurrent Neural Networks (LRNNs) such as Mamba, RWKV, GLA, mLSTM, and DeltaNet have emerged as efficient alternatives to Transformers for long sequences. However, both Tran…
Is Mamba Capable of In-Context Learning?
Riccardo Grazzi, Julien Siems, Simon Schrodi +2
State of the art foundation models such as GPT-4 perform surprisingly well at in-context learning (ICL), a variant of meta-learning concerning the learned ability to solve tasks du…
Learning invariant representations of time-homogeneous stochastic dynamical systems
Vladimir R. Kostic, Pietro Novelli, Riccardo Grazzi +2
We consider the general class of time-homogeneous stochastic dynamical systems, both discrete and continuous, and study the problem of learning a representation of the state that f…
Learning-to-Learn Stochastic Gradient Descent with Biased Regularization
Giulia Denevi, Carlo Ciliberto, Riccardo Grazzi +1
We study the problem of learning-to-learn: inferring a learning algorithm that works well on tasks sampled from an unknown distribution. As class of algorithms we consider Stochast…