activity
20172020
most citedVariants of RMSProp and Adagrad with Logarithmic Regret Bounds

110 citations · 115 across the 3 of their papers we have counts for

collaborators

6 papers

math.OC2020

Global Convergence of Model Function Based Bregman Proximal Minimization Algorithms

Mahesh Chandra Mukkamala, Jalal Fadili, Peter Ochs

Lipschitz continuity of the gradient mapping of a continuously differentiable function plays a crucial role in designing various optimization algorithms. However, many functions ar…

math.OC20195 cited

Bregman Proximal Framework for Deep Linear Neural Networks

Mahesh Chandra Mukkamala, Felix Westerkamp, Emanuel Laude +2

A typical assumption for the analysis of first order optimization methods is the Lipschitz continuity of the gradient of the objective function. However, for many practical applica…

math.OC2019

Beyond Alternating Updates for Matrix Factorization with Inertial Bregman Proximal Gradient Algorithms

Mahesh Chandra Mukkamala, Peter Ochs

Matrix Factorization is a popular non-convex optimization problem, for which alternating minimization schemes are mostly used. They usually suffer from the major drawback that the…

cs.LG2018

On the loss landscape of a class of deep neural networks with no bad local valleys

Quynh Nguyen, Mahesh Chandra Mukkamala, Matthias Hein

We identify a class of over-parameterized deep neural networks with standard activation functions and cross-entropy loss which provably have no bad local valley, in the sense that…

cs.LG2018

Neural Networks Should Be Wide Enough to Learn Disconnected Decision Regions

Quynh Nguyen, Mahesh Chandra Mukkamala, Matthias Hein

In the recent literature the important role of depth in deep learning has been emphasized. In this paper we argue that sufficient width of a feedforward network is equally importan…

cs.LG2017110 cited

Variants of RMSProp and Adagrad with Logarithmic Regret Bounds

Mahesh Chandra Mukkamala, Matthias Hein

Adaptive gradient methods have become recently very popular, in particular as they have been shown to be useful in the training of deep neural networks. In this paper we have analy…