5 papers
Smoothing DiLoCo with Primal Averaging for Faster Training of LLMs
Aaron Defazio, Konstantin Mishchenko, Parameswaran Raman +2
We propose Generalized Primal Averaging (GPA), an extension of Nesterov's method that unifies and generalizes recent averaging-based optimizers like single-worker DiLoCo and Schedu…
Hardware-Aware Parallel Prompt Decoding for Memory-Efficient Acceleration of LLM Inference
Hao Mark Chen, Wayne Luk, Ka Fai Cedric Yiu +4
The auto-regressive decoding of Large Language Models (LLMs) results in significant overheads in their hardware performance. While recent research has investigated various speculat…
Analysis of an Idealized Stochastic Polyak Method and its Application to Black-Box Model Distillation
Robert M. Gower, Guillaume Garrigos, Nicolas Loizou +3
We provide a general convergence theorem of an idealized stochastic Polyak step size called SPS. Besides convexity, we only assume a local expected gradient bound, that include…
The Road Less Scheduled
Aaron Defazio, Xingyu Alice Yang, Harsh Mehta +3
Existing learning rate schedules that do not require specification of the optimization stopping step T are greatly out-performed by learning rate schedules that depend on T. We pro…
Optimal Linear Decay Learning Rate Schedules and Further Refinements
Aaron Defazio, Ashok Cutkosky, Harsh Mehta +1
Learning rate schedules used in practice bear little resemblance to those recommended by theory. We close much of this theory/practice gap, and as a consequence are able to derive…