6 papers · 1 filter
Extracting Dual Solutions via Primal Optimizers
Yair Carmon, Arun Jambulapati, Liam O'Carroll +1
We provide a general method to convert a "primal" black-box algorithm for solving regularized convex-concave minimax optimization problems into an algorithm for solving the associa…
Accelerated Parameter-Free Stochastic Optimization
Itai Kreisler, Maor Ivgi, Oliver Hinder +1
We propose a method that achieves near-optimal rates for smooth stochastic convex optimization and requires essentially no prior knowledge of problem parameters. This improves on p…
Malign Overfitting: Interpolation Can Provably Preclude Invariance
Yoav Wald, Gal Yona, Uri Shalit +1
Learned classifiers should often possess certain invariance properties meant to encourage fairness, robustness, or out-of-distribution generalization. However, multiple recent work…
The Price of Adaptivity in Stochastic Convex Optimization
Yair Carmon, Oliver Hinder
We prove impossibility results for adaptivity in non-smooth stochastic convex optimization. Given a set of problem parameters we wish to adapt to, we define a "price of adaptivity"…
Language models scale reliably with over-training and on downstream tasks
Samir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar +22
Scaling laws are useful guides for derisking expensive training runs, as they predict performance of large models using cheaper, small-scale experiments. However, there remain gaps…
Making SGD Parameter-Free
Yair Carmon, Oliver Hinder
We develop an algorithm for parameter-free stochastic convex optimization (SCO) whose rate of convergence is only a double-logarithmic factor larger than the optimal rate for the c…