3 papers
cs.LG2026
Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases
Jingwen Liu, Ezra Edelman, Surbhi Goel +1
This work investigates the ``small-vs-large gap'', where repeating on fewer samples can lead to compute saving during training compared to using a larger dataset. This is observed…
cs.LG2026
The two clocks and the innovation window: When and how generative models learn rules
Binxu Wang, Emma Lucia Byrnes Finn, Bingbin Liu
Generative models trained on finite data face a fundamental tension: their score-matching or next-token objective converges to the empirical training distribution rather than the p…
cs.LG2025
Adam or Gauss-Newton? A Comparative Study In Terms of Basis Alignment and SGD Noise
Bingbin Liu, Rachit Bansal, Depen Morwani +3
Diagonal preconditioners are computationally feasible approximate to second-order optimizers, which have shown significant promise in accelerating training of deep learning models.…