256 citations · 457 across the 8 of their papers we have counts for
4 papers · 1 filter
BYOL works even without batch statistics
Pierre H. Richemond, Jean-Bastien Grill, Florent Altché +8
Bootstrap Your Own Latent (BYOL) is a self-supervised learning approach for image representation. From an augmented view of an image, BYOL trains an online network to predict a tar…
On the Generalization Benefit of Noise in Stochastic Gradient Descent
Samuel L. Smith, Erich Elsen, Soham De
It has long been argued that minibatch stochastic gradient descent can generalize better than large batch gradient descent in deep neural networks. However recent papers have quest…
Modeling Citation Trajectories of Scientific Papers
Dattatreya Mohapatra, Siddharth Pal, Soham De +2
Several network growth models have been proposed in the literature that attempt to incorporate properties of citation networks. Generally, these models aim at retaining the degree…
Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep Networks
Soham De, Samuel L. Smith
Batch normalization dramatically increases the largest trainable depth of residual networks, and this benefit has been crucial to the empirical success of deep residual networks on…