activity
20182026
most citedResurrecting Recurrent Neural Networks for Long Sequences

43 citations · 114 across the 40 of their papers we have counts for

collaborators
Showing cs.LGShow all

27 papers · 1 filter

cs.LG2026

Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

Omar Coser, Antonio Orvieto, Paolo Soda +1

Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal…

cs.LG2026

Muown Implicitly Performs Angular Step-size Decay

Florian Hübler, Kai Lion, Antonio Orvieto +1

Matrix-aware optimizers such as Muon and Muown have recently shown strong empirical performance for pre-training Transformers. In particular, Muown separates each weight matrix int…

cs.LG2026

Towards Understanding Self-Pretraining for Sequence Classification

Omar Coser, Loredana Zollo, Paolo Soda +1

Amos et al. (2024) showed that the accuracy of Transformer models in sequence classification can be significantly improved by first pretraining with a masked token prediction objec…

cs.LG2026

Explaining Grokking in Transformers through the Lens of Inductive Bias

Jaisidh Singh, Diganta Misra, Antonio Orvieto

We investigate grokking in transformers through the lens of inductive bias: dispositions arising from architecture or optimization that let the network prefer one solution over ano…

cs.LG2026

Universal Dynamics of Warmup Stable Decay: understanding WSD beyond Transformers

Annalisa Belloni, Lorenzo Noci, Antonio Orvieto

The Warmup Stable Decay (WSD) learning rate scheduler has recently become popular, largely due to its good performance and flexibility when training large language models. It remai…

cs.LG2025

Design Principles for Sequence Models via Coefficient Dynamics

Jerome Sieber, Antonio Orvieto, Melanie N. Zeilinger +1

Deep sequence models, ranging from Transformers and State Space Models (SSMs) to more recent approaches such as gated linear RNNs, fundamentally compute outputs as linear combinati…