160 citations · 261 across the 7 of their papers we have counts for
11 papers
On the Adversarial Robustness of Mixture of Experts
Joan Puigcerver, Rodolphe Jenatton, Carlos Riquelme +2
Adversarial robustness is a key desirable property of neural networks. It has been empirically shown to be affected by their sizes, with larger networks being typically more robust…
Learning to Merge Tokens in Vision Transformers
Cedric Renggli, André Susano Pinto, Neil Houlsby +3
Transformers are widely applied to solve natural language understanding and computer vision tasks. While scaling up these architectures leads to improved performance, it often come…
Scaling Vision with Sparse Mixture of Experts
Carlos Riquelme, Joan Puigcerver, Basil Mustafa +5
Sparsely-gated Mixture of Experts networks (MoEs) have demonstrated excellent scalability in Natural Language Processing. In Computer Vision, however, almost all performant network…
Deep Ensembles for Low-Data Transfer Learning
Basil Mustafa, Carlos Riquelme, Joan Puigcerver +3
In the low-data regime, it is difficult to train good supervised models from scratch. Instead practitioners turn to pre-trained models, leveraging transfer learning. Ensembling is…
Scalable Transfer Learning with Expert Models
Joan Puigcerver, Carlos Riquelme, Basil Mustafa +5
Transfer of pre-trained representations can improve sample efficiency and reduce computational requirements for new tasks. However, representations used for transfer are usually ge…
On Last-Layer Algorithms for Classification: Decoupling Representation from Uncertainty Estimation
Nicolas Brosse, Carlos Riquelme, Alice Martin +2
Uncertainty quantification for deep learning is a challenging open problem. Bayesian statistics offer a mathematically grounded framework to reason about uncertainties; however, ap…