activity
20112017
most citedPractical Gauss-Newton Optimisation for Deep Learning

35 citations · 70 across the 5 of their papers we have counts for

collaborators

5 papers

stat.ML201721 cited

Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning

Zhen He, Shaobing Gao, Liang Xiao +3

Long Short-Term Memory (LSTM) is a popular approach to boosting the ability of Recurrent Neural Networks to store longer term temporal information. The capacity of an LSTM network…

stat.ML201735 cited

Practical Gauss-Newton Optimisation for Deep Learning

Aleksandar Botev, Hippolyt Ritter, David Barber

We present an efficient block-diagonal ap- proximation to the Gauss-Newton matrix for feedforward neural networks. Our result- ing algorithm is competitive against state- of-the-ar…

stat.ML20161 cited

Nesterov's Accelerated Gradient and Momentum as approximations to Regularised Update Descent

Aleksandar Botev, Guy Lever, David Barber

We present a unifying framework for adapting the update direction in gradient-based iterative optimization methods. As natural special cases we re-derive classical momentum and Nes…

stat.ML20162 cited

Dealing with a large number of classes -- Likelihood, Discrimination or Ranking?

David Barber, Aleksandar Botev

We consider training probabilistic classifiers in the case of a large number of classes. The number of classes is assumed too large to perform exact normalisation over all classes.…

cs.AI201111 cited

Variational Cumulant Expansions for Intractable Distributions

D. Barber, P. de van Laar

Intractable distributions present a common difficulty in inference within the probabilistic knowledge representation framework and variational methods have recently been popular in…