collaborators

13 papers

cs.LG2026

Is Self-Pretraining really useful to improve diagnosis in medical Time Series?

Omar Coser, Antonio Orvieto, Paolo Soda +1

Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal…

cs.LG2026

Muown Implicitly Performs Angular Step-size Decay

Florian Hübler, Kai Lion, Antonio Orvieto +1

Matrix-aware optimizers such as Muon and Muown have recently shown strong empirical performance for pre-training Transformers. In particular, Muown separates each weight matrix int…

cs.AI2026

Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers

Sajad Movahedi, Vera Milovanović, Shlomo Libo Feigin +5

Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by loo…

cs.LG2026

Towards Understanding Self-Pretraining for Sequence Classification

Omar Coser, Loredana Zollo, Paolo Soda +1

Amos et al. (2024) showed that the accuracy of Transformer models in sequence classification can be significantly improved by first pretraining with a masked token prediction objec…

cs.LG2026

Design Principles for Sequence Models via Coefficient Dynamics

Jerome Sieber, Antonio Orvieto, Melanie N. Zeilinger +1

Deep sequence models, ranging from Transformers and State Space Models (SSMs) to more recent approaches such as gated linear RNNs, fundamentally compute outputs as linear combinati…

cs.LG2026

Explaining Grokking in Transformers through the Lens of Inductive Bias

Jaisidh Singh, Diganta Misra, Antonio Orvieto

We investigate grokking in transformers through the lens of inductive bias: dispositions arising from architecture or optimization that let the network prefer one solution over ano…