activity
20182026
most citedLearning-to-Learn Stochastic Gradient Descent with Biased Regularization

45 citations · 74 across the 6 of their papers we have counts for

collaborators
Showing cs.LGShow all

6 papers · 1 filter

cs.LG2026

Learning State-Tracking from Code Using Linear RNNs

Julien Siems, Riccardo Grazzi, Korbinian Pöppel +3

Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence models architectures like Transformers a…

cs.LG2025

DeltaProduct: Improving State-Tracking in Linear RNNs via Householder Products

Julien Siems, Timur Carstensen, Arber Zela +3

Linear Recurrent Neural Networks (linear RNNs) have emerged as competitive alternatives to Transformers for sequence modeling, offering efficient training and linear-time inference…

cs.LG20241 cited

Unlocking State-Tracking in Linear RNNs Through Negative Eigenvalues

Riccardo Grazzi, Julien Siems, Arber Zela +3

Linear Recurrent Neural Networks (LRNNs) such as Mamba, RWKV, GLA, mLSTM, and DeltaNet have emerged as efficient alternatives to Transformers for long sequences. However, both Tran…

cs.LG20245 cited

Is Mamba Capable of In-Context Learning?

Riccardo Grazzi, Julien Siems, Simon Schrodi +2

State of the art foundation models such as GPT-4 perform surprisingly well at in-context learning (ICL), a variant of meta-learning concerning the learned ability to solve tasks du…

cs.LG2023

Learning invariant representations of time-homogeneous stochastic dynamical systems

Vladimir R. Kostic, Pietro Novelli, Riccardo Grazzi +2

We consider the general class of time-homogeneous stochastic dynamical systems, both discrete and continuous, and study the problem of learning a representation of the state that f…

cs.LG201945 cited

Learning-to-Learn Stochastic Gradient Descent with Biased Regularization

Giulia Denevi, Carlo Ciliberto, Riccardo Grazzi +1

We study the problem of learning-to-learn: inferring a learning algorithm that works well on tasks sampled from an unknown distribution. As class of algorithms we consider Stochast…