10 citations · 10 across the 4 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Is Self-Pretraining really useful to improve diagnosis in medical Time Series?
Omar Coser, Antonio Orvieto, Paolo Soda +1
Inspired by recent evidence that transformer architectures benefit from Self-PreTraining (SPT) on long-context benchmarks, we investigate whether similar gains extend to multimodal…
cs.LG2026
Towards Understanding Self-Pretraining for Sequence Classification
Omar Coser, Loredana Zollo, Paolo Soda +1
Amos et al. (2024) showed that the accuracy of Transformer models in sequence classification can be significantly improved by first pretraining with a masked token prediction objec…