35 citations · 70 across the 5 of their papers we have counts for
5 papers
Wider and Deeper, Cheaper and Faster: Tensorized LSTMs for Sequence Learning
Zhen He, Shaobing Gao, Liang Xiao +3
Long Short-Term Memory (LSTM) is a popular approach to boosting the ability of Recurrent Neural Networks to store longer term temporal information. The capacity of an LSTM network…
Practical Gauss-Newton Optimisation for Deep Learning
Aleksandar Botev, Hippolyt Ritter, David Barber
We present an efficient block-diagonal ap- proximation to the Gauss-Newton matrix for feedforward neural networks. Our result- ing algorithm is competitive against state- of-the-ar…
Nesterov's Accelerated Gradient and Momentum as approximations to Regularised Update Descent
Aleksandar Botev, Guy Lever, David Barber
We present a unifying framework for adapting the update direction in gradient-based iterative optimization methods. As natural special cases we re-derive classical momentum and Nes…
Dealing with a large number of classes -- Likelihood, Discrimination or Ranking?
David Barber, Aleksandar Botev
We consider training probabilistic classifiers in the case of a large number of classes. The number of classes is assumed too large to perform exact normalisation over all classes.…
Variational Cumulant Expansions for Intractable Distributions
D. Barber, P. de van Laar
Intractable distributions present a common difficulty in inference within the probabilistic knowledge representation framework and variational methods have recently been popular in…