Showing cs.LGShow all
3 papers · 1 filter
cs.LG2025
Continual Pre-training of MoEs: How robust is your router?
Benjamin Thérien, Charles-Étienne Joseph, Zain Sarwar +7
Sparsely-activated Mixture of Experts (MoE) transformers are promising architectures for foundation models. Compared to dense transformers that require the same amount of floating-…
cs.LG2025
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
Vaibhav Singh, Paul Janson, Paria Mehrbod +4
The ever-growing availability of unlabeled data presents both opportunities and challenges for training artificial intelligence systems. While self-supervised learning (SSL) has em…
cs.LG2022
A Closer Look at Robustness to L-infinity and Spatial Perturbations and their Composition
Luke Rowe, Benjamin Thérien, Krzysztof Czarnecki +1
In adversarial machine learning, the popular threat model has been the focus of much previous work. While this mathematical definition of imperceptibility successfull…