Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States
Luca Herranz-Celotti, Vincent Guigue
Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor (a matrix), limiting the order of interactio…
cs.LG2023
Stabilizing RNN Gradients through Pre-training
Luca Herranz-Celotti, Jean Rouat
Numerous theories of learning propose to prevent the gradient from exponential growth with depth or time, to stabilize and improve training. Typically, these analyses are conducted…