284 citations · 405 across the 19 of their papers we have counts for
5 papers · 1 filter
Phenomenology of Double Descent in Finite-Width Neural Networks
Sidak Pal Singh, Aurelien Lucchi, Thomas Hofmann +1
`Double descent' delineates the generalization behaviour of models depending on the regime they belong to: under- or over-parameterized. The current theoretical understanding behin…
Batch Normalization Provably Avoids Rank Collapse for Randomly Initialised Deep Networks
Hadi Daneshmand, Jonas Kohler, Francis Bach +2
Randomly initialized neural networks are known to become harder to train with increasing depth, unless architectural enhancements like residual connections and batch normalization…
Adversarially Robust Training through Structured Gradient Regularization
Kevin Roth, Aurelien Lucchi, Sebastian Nowozin +1
We propose a novel data-dependent structured gradient regularizer to increase the robustness of neural networks vis-a-vis adversarial perturbations. Our regularizer can be derived…
Exponential convergence rates for Batch Normalization: The power of length-direction decoupling in non-convex optimization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi +3
Normalization techniques such as Batch Normalization have been applied successfully for training deep neural networks. Yet, despite its apparent empirical benefits, the reasons beh…
Generator Reversal
Yannic Kilcher, Aurélien Lucchi, Thomas Hofmann
We consider the problem of training generative models with deep neural networks as generators, i.e. to map latent codes to data points. Whereas the dominant paradigm combines simpl…