collaborators

7 papers

cs.LG2026

How to Allocate Your Tokens? Scaling Laws with Training Steps and Batch Size

Fabian Schaipp

We propose a scaling law that takes into account model size and training data while explicitly splitting the latter into training steps and batch size (called three-term law). Fitt…

math.OC2026

Step-Size Stability in Stochastic Optimization: A Theoretical Perspective

Fabian Schaipp, Robert M. Gower, Adrien Taylor

We present a theoretical analysis of stochastic optimization methods in terms of their sensitivity with respect to the step size. We identify a key quantity that, for each method,…

q-bio.QM2026

Sparse regression, classification, and microbial network estimation in QIIME2 with q2-classo and q2-gglasso

Oleg Vlasovets, Fabian Schaipp, Leo Simpson +3

Motivation: Statistical analysis of microbial count data derived from 16S rRNA or metagenomics sequencing poses unique challenges due to the sparse, compositional, and high-dimensi…

stat.ML2025

Tracking the Median of Gradients with a Stochastic Proximal Point Method

Fabian Schaipp, Guillaume Garrigos, Umut Simsekli +1

There are several applications of stochastic optimization where one can benefit from a robust estimate of the gradient. For example, domains such as distributed learning with corru…

cs.LG2025

Optimization Benchmark for Diffusion Models on Dynamical Systems

Fabian Schaipp

The training of diffusion models is often absent in the evaluation of new optimization techniques. In this work, we benchmark recent optimization algorithms for training a diffusio…

cs.LG2025

The Surprising Agreement Between Convex Optimization Theory and Learning-Rate Scheduling for Large Model Training

Fabian Schaipp, Alexander Hägele, Adrien Taylor +2

We show that learning-rate schedules for large model training behave surprisingly similar to a performance bound from non-smooth convex optimization theory. We provide a bound for…