activity
20242026
collaborators

9 papers

cs.LG2026

A projection-based framework for gradient-free and parallel learning

Andreas Bergmeister, Manish Krishan Lal, Stefanie Jegelka +1

We present a feasibility-seeking approach to neural network training. This mathematical optimization framework is distinct from conventional gradient-based loss minimization and us…

cs.LG2025

Near-Optimal Algorithms for Group Distributionally Robust Optimization and Beyond

Tasuku Soma, Khashayar Gatmiry, Sharut Gupta +1

Distributionally robust optimization (DRO) can improve the robustness and fairness of learning methods. In this paper, we devise stochastic algorithms for a class of DRO problems i…

cs.LG2024

On the Role of Depth and Looping for In-Context Learning with Task Diversity

Khashayar Gatmiry, Nikunj Saunshi, Sashank J. Reddi +2

The intriguing in-context learning (ICL) abilities of deep Transformer models have lately garnered significant attention. By studying in-context linear regression on unimodal Gauss…

cs.LG2024

Computing Optimal Regularizers for Online Linear Optimization

Khashayar Gatmiry, Jon Schneider, Stefanie Jegelka

Follow-the-Regularized-Leader (FTRL) algorithms are a popular class of learning algorithms for online linear optimization (OLO) that guarantee sub-linear regret, but the choice of…

cs.LG2024

Simplicity Bias via Global Convergence of Sharpness Minimization

Khashayar Gatmiry, Zhiyuan Li, Sashank J. Reddi +1

The remarkable generalization ability of neural networks is usually attributed to the implicit bias of SGD, which often yields models with lower complexity using simpler (e.g. line…

cs.LG2024

Can Looped Transformers Learn to Implement Multi-step Gradient Descent for In-context Learning?

Khashayar Gatmiry, Nikunj Saunshi, Sashank J. Reddi +2

The remarkable capability of Transformers to do reasoning and few-shot learning, without any fine-tuning, is widely conjectured to stem from their ability to implicitly simulate a…