4 papers
Continual Pre-training of MoEs: How robust is your router?
Benjamin Thérien, Charles-Étienne Joseph, Zain Sarwar +7
Sparsely-activated Mixture of Experts (MoE) transformers are promising architectures for foundation models. Compared to dense transformers that require the same amount of floating-…
Beyond Cosine Decay: On the effectiveness of Infinite Learning Rate Schedule for Continual Pre-training
Vaibhav Singh, Paul Janson, Paria Mehrbod +4
The ever-growing availability of unlabeled data presents both opportunities and challenges for training artificial intelligence systems. While self-supervised learning (SSL) has em…
A Closer Look at Robustness to L-infinity and Spatial Perturbations and their Composition
Luke Rowe, Benjamin Thérien, Krzysztof Czarnecki +1
In adversarial machine learning, the popular threat model has been the focus of much previous work. While this mathematical definition of imperceptibility successfull…
Interpretable Deep Tracking
Benjamin Thérien, Krzysztof Czarnecki
Imagine experiencing a crash as the passenger of an autonomous vehicle. Wouldn't you want to know why it happened? Current end-to-end optimizable deep neural networks (DNNs) in 3D…