264 citations · 713 across the 13 of their papers we have counts for
4 papers · 1 filter
BYOL works even without batch statistics
Pierre H. Richemond, Jean-Bastien Grill, Florent Altché +8
Bootstrap Your Own Latent (BYOL) is a self-supervised learning approach for image representation. From an augmented view of an image, BYOL trains an online network to predict a tar…
Cold Posteriors and Aleatoric Uncertainty
Ben Adlam, Jasper Snoek, Samuel L. Smith
Recent work has observed that one can outperform exact inference in Bayesian neural networks by tuning the "temperature" of the posterior on a validation set (the "cold posterior"…
On the Generalization Benefit of Noise in Stochastic Gradient Descent
Samuel L. Smith, Erich Elsen, Soham De
It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have quest…
Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
Soham De, Samuel L. Smith
Batch normalization dramatically increases the largest trainable depth of residual networks, and this benefit has been crucial to the empirical success of deep residual networks on…