1 paper
Fabian Schaipp, Alexander Hägele, Adrien Taylor +2
We show that learning-rate schedules for large model training behave surprisingly similar to a performance bound from non-smooth convex optimization theory. We provide a bound for…