3 papers
cs.LG2026
RunningTensor: Generalizing Linear Attention to Higher-Order Recurrent States
Luca Herranz-Celotti, Vincent Guigue
Linear attention and state-space models provide linear-time sequence modeling, but their recurrent memory remains a second-order tensor (a matrix), limiting the order of interactio…
cs.LG2023
Stabilizing RNN Gradients through Pre-training
Luca Herranz-Celotti, Jean Rouat
Numerous theories of learning propose to prevent the gradient from exponential growth with depth or time, to stabilize and improve training. Typically, these analyses are conducted…
cs.CL2023
Less is More! A slim architecture for optimal language translation
Luca Herranz-Celotti, Ermal Rrapaj
The softmax attention mechanism has emerged as a noteworthy development in the field of Artificial Intelligence research, building on the successes of Transformer-based architectur…