4 papers
Reparameterizing Mirror Descent as Gradient Descent
Ehsan Amid, Manfred K. Warmuth
Most of the recent successful applications of neural networks have been based on training with gradient descent updates. However, for some small networks, other mirror descent upda…
An Implicit Form of Krasulina's k-PCA Update without the Orthonormality Constraint
Ehsan Amid, Manfred K. Warmuth
We shed new insights on the two commonly used updates for the online -PCA problem, namely, Krasulina's and Oja's updates. We show that Krasulina's update corresponds to a projec…
Robust Bi-Tempered Logistic Loss Based on Bregman Divergences
Ehsan Amid, Manfred K. Warmuth, Rohan Anil +1
We introduce a temperature into the exponential function and replace the softmax output layer of neural nets by a high temperature generalization. Similarly, the logarithm in the l…
Divergence-Based Motivation for Online EM and Combining Hidden Variable Models
Ehsan Amid, Manfred K. Warmuth
Expectation-Maximization (EM) is a prominent approach for parameter estimation of hidden (aka latent) variable models. Given the full batch of data, EM forms an upper-bound of the…