8 papers
One-shot Conditional Sampling: MMD meets Nearest Neighbors
Anirban Chatterjee, Sayantan Choudhury, Rohan Hore
How can we generate samples from a conditional distribution that we never fully observe? This question arises across a broad range of applications in both modern machine learning a…
Gradient Clipping Beyond Vector Norms: A Spectral Approach for Matrix-Valued Parameters
Alexander Yukhimchuk, Mladen Kolar, Martin TakÃ¡Ä +1
Gradient clipping is a standard safeguard for training neural networks under noisy, heavy-tailed stochastic gradients; yet, most clipping rules treat all parameters as vectors and…
Muon with Nesterov Momentum: Heavy-Tailed Noise and (Randomized) Inexact Polar Decomposition
Sayantan Choudhury, Xiaoran Cheng, Martin TakÃ¡Ä +2
Most first-order optimizers treat matrix-valued parameters as vectors, ignoring the intrinsic geometry of hidden-layer weights in neural networks. Muon addresses this mismatch by u…
Doubly-Unlinked Regression for Dependent Data
Anik Burman, Sayantan Choudhury, Debangan Dey
Shuffled regression concerns settings in which covariates and responses are observed without their correct pairing. In dependent-data problems, a second form of missing corresponde…
Multiplayer Federated Learning: Reaching Equilibrium with Less Communication
TaeHo Yoon, Sayantan Choudhury, Nicolas Loizou
Traditional Federated Learning (FL) approaches assume collaborative clients with aligned objectives working towards a shared global model. However, in many real-world scenarios, cl…
Extragradient Method for -Lipschitz Root-finding Problems
Sayantan Choudhury, Nicolas Loizou
Introduced by Korpelevich in 1976, the extragradient method (EG) has become a cornerstone technique for solving min-max optimization, root-finding problems, and variational inequal…