7 papers
How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size
Fabian Schaipp
We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law). Fitt…
Step-Size Stability in Stochastic Optimization: A Theoretical Perspective
Fabian Schaipp, Robert M. Gower, Adrien Taylor
We present a theoretical analysis of stochastic optimization methods in terms of their sensitivity with respect to the step size. We identify a key quantity that, for each method,…
Sparse regression, classification, and microbial network estimation in QIIME2 with q2-classo and q2-gglasso
Oleg Vlasovets, Fabian Schaipp, Leo Simpson +3
Motivation: Statistical analysis of microbial count data derived from 16S rRNA or metagenomics sequencing poses unique challenges due to the sparse, compositional, and high-dimensi…
Tracking the Median of Gradients with a Stochastic Proximal Point Method
Fabian Schaipp, Guillaume Garrigos, Umut Simsekli +1
There are several applications of stochastic optimization where one can benefit from a robust estimate of the gradient. For example, domains such as distributed learning with corru…
Optimization Benchmark for Diffusion Models on Dynamical Systems
Fabian Schaipp
The training of diffusion models is often absent in the evaluation of new optimization techniques. In this work, we benchmark recent optimization algorithms for training a diffusio…
The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training
Fabian Schaipp, Alexander Hägele, Adrien Taylor +2
We show that learning-rate schedules for large model training behave surprisingly similar to a performance bound from non-smooth convex optimization theory. We provide a bound for…