13 papers
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
Omar Coser, Antonio Orvieto, Paolo Soda +1
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal…
Muown Implicitly Performs Angular Step-size Decay
Florian Hübler, Kai Lion, Antonio Orvieto +1
Matrix-aware optimizers such as Muon and Muown have recently shown strong empirical performance for pre-training Transformers. In particular, Muown separates each weight matrix int…
Fixed-Point Reasoners: Stable and Adaptive Deep Looped Transformers
Sajad Movahedi, Vera MilovanoviÄ, Shlomo Libo Feigin +5
Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning. The number of effective layers reached by loo…
Towards Understanding Self-Pretraining for Sequence Classification
Omar Coser, Loredana Zollo, Paolo Soda +1
Amos et al. (2024) showed that the accuracy of Transformer models in sequence classification can be significantly improved by first pretraining with a masked token prediction objec…
Design Principles for Sequence Models via Coefficient Dynamics
Jerome Sieber, Antonio Orvieto, Melanie N. Zeilinger +1
Deep sequence models, ranging from Transformers and State Space Models (SSMs) to more recent approaches such as gated linear RNNs, fundamentally compute outputs as linear combinati…
Explaining Grokking in Transformers through the Lens of Inductive Bias
Jaisidh Singh, Diganta Misra, Antonio Orvieto
We investigate grokking in transformers through the lens of inductive bias: dispositions arising from architecture or optimization that let the network prefer one solution over ano…