4 papers
Provable Benefit of Sign Descent: A Minimal Model Under Heavy-Tailed Class Imbalance
Robin Yadav, Shuo Xie, Tianhao Wang +1
Adaptive optimization methods (such as Adam) play a major role in LLM pretraining, significantly outperforming Gradient Descent (GD). Recent studies have proposed new smoothness as…
RETRO SYNFLOW: Discrete Flow Matching for Accurate and Diverse Single-Step Retrosynthesis
Robin Yadav, Qi Yan, Guy Wolf +2
A fundamental problem in organic chemistry is identifying and predicting the series of reactions that synthesize a desired target product molecule. Due to the combinatorial nature…
Local Curvature Descent: Squeezing More Curvature out of Standard and Polyak Gradient Descent
Peter Richtárik, Simone Maria Giancola, Dymitr Lubczyk +1
We contribute to the growing body of knowledge on more powerful and adaptive stepsizes for convex optimization, empowered by local curvature information. We do not go the route of…
Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Frederik Kunstner, Robin Yadav, Alan Milligan +2
Adam has been shown to outperform gradient descent on large language models by a larger margin than on other tasks, but it is unclear why. We show that a key factor in this perform…