6 papers · 1 filter
Exponentiated Gradient Reweighting for Robust Training Under Label Noise and Beyond
Negin Majidi, Ehsan Amid, Hossein Talebi +1
Many learning tasks in machine learning can be viewed as taking a gradient step towards minimizing the average loss of a batch of examples in each training iteration. When noise is…
A case where a spindly two-layer linear network whips any neural network with a fully connected input layer
Manfred K. Warmuth, Wojciech Kotłowski, Ehsan Amid
It was conjectured that any neural network of any structure and arbitrary differentiable transfer functions at the nodes cannot learn the following problem sample efficiently when…
Reparameterizing Mirror Descent as Gradient Descent
Ehsan Amid, Manfred K. Warmuth
Most of the recent successful applications of neural networks have been based on training with gradient descent updates. However, for some small networks, other mirror descent upda…
An Implicit Form of Krasulina's k-PCA Update without the Orthonormality Constraint
Ehsan Amid, Manfred K. Warmuth
We shed new insights on the two commonly used updates for the online -PCA problem, namely, Krasulina's and Oja's updates. We show that Krasulina's update corresponds to a projec…
Robust Bi-Tempered Logistic Loss Based on Bregman Divergences
Ehsan Amid, Manfred K. Warmuth, Rohan Anil +1
We introduce a temperature into the exponential function and replace the softmax output layer of neural nets by a high temperature generalization. Similarly, the logarithm in the l…
Divergence-Based Motivation for Online EM and Combining Hidden Variable Models
Ehsan Amid, Manfred K. Warmuth
Expectation-Maximization (EM) is a prominent approach for parameter estimation of hidden (aka latent) variable models. Given the full batch of data, EM forms an upper-bound of the…