collaborators

10 papers

cs.LG2026

RESIST: Resilient Decentralized Learning Using Consensus Gradient Descent

Cheng Fang, Rishabh Dixit, Waheed U. Bajwa +1

Empirical risk minimization (ERM) is a cornerstone of modern machine learning (ML), supported by advances in optimization theory that ensure efficient solutions with provable algor…

math.OC2026

Accelerated Gradient Methods for Nonconvex Optimization: Escape Trajectories From Strict Saddle Points and Convergence to Local Minima

Rishabh Dixit, Mert Gurbuzbalaban, Waheed U. Bajwa

This paper considers the problem of understanding the behavior of a general class of accelerated gradient methods on smooth nonconvex functions. Motivated by some recent works that…

math.OC2026

Accelerated Gradient Methods with Biased Gradient Estimates: Risk Sensitivity, High-Probability Guarantees, and Large Deviation Bounds

Mert Gürbüzbalaban, Yasa Syed, Necdet Serhat Aybat

We study trade-offs between convergence rate and robustness to gradient errors in the context of first-order methods. Our focus is on generalized momentum methods (GMMs)--a broad c…

stat.ML2025

Rényi Differential Privacy for Heavy-Tailed SDEs via Fractional Poincaré Inequalities

Benjamin Dupuis, Mert Gürbüzbalaban, Umut Şimşekli +3

Characterizing the differential privacy (DP) of learning algorithms has become a major challenge in recent years. In parallel, many studies suggested investigating the behavior of…

math.OC2025

DIGing--SGLD: Decentralized and Scalable Langevin Sampling over Time--Varying Networks

Waheed U. Bajwa, Mert Gurbuzbalaban, Mustafa Ali Kutbay +2

Sampling from a target distribution induced by training data is central to Bayesian learning, with Stochastic Gradient Langevin Dynamics (SGLD) serving as a key tool for scalable p…

cs.LG2025

Generalized EXTRA stochastic gradient Langevin dynamics

Mert Gurbuzbalaban, Mohammad Rafiqul Islam, Xiaoyu Wang +1

Langevin algorithms are popular Markov Chain Monte Carlo methods for Bayesian learning, particularly when the aim is to sample from the posterior distribution of a parametric model…