19 citations · 19 across the 1 of their papers we have counts for
5 papers
This Looks Like That... Does it? Shortcomings of Latent Space Prototype Interpretability in Deep Networks
Adrian Hoffmann, Claudio Fanconi, Rahul Rade +1
Deep neural networks that yield human interpretable decisions by architectural design have lately become an increasingly popular alternative to post hoc interpretation of tradition…
Batch Normalization Provably Avoids Rank Collapse for Randomly Initialised Deep Networks
Hadi Daneshmand, Jonas Kohler, Francis Bach +2
Randomly initialized neural networks are known to become harder to train with increasing depth, unless architectural enhancements like residual connections and batch normalization…
Exponential convergence rates for Batch Normalization: The power of length-direction decoupling in non-convex optimization
Jonas Kohler, Hadi Daneshmand, Aurelien Lucchi +3
Normalization techniques such as Batch Normalization have been applied successfully for training deep neural networks. Yet, despite its apparent empirical benefits, the reasons beh…
Escaping Saddles with Stochastic Gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi +1
We analyze the variance of stochastic gradients along negative curvature directions in certain non-convex machine learning models and show that stochastic gradients exhibit a stron…
Sub-sampled Cubic Regularization for Non-convex Optimization
Jonas Moritz Kohler, Aurelien Lucchi
We consider the minimization of non-convex functions that typically arise in machine learning. Specifically, we focus our attention on a variant of trust region methods known as cu…