3 papers
cs.LG2026
Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases
Jingwen Liu, Ezra Edelman, Surbhi Goel +1
This work investigates the ``small-vs-large gap'', where repeating on fewer samples can lead to compute saving during training compared to using a larger dataset. This is observed…
cs.LG2026
Reliable Abstention under Adversarial Injections: Tight Lower Bounds and New Upper Bounds
Ezra Edelman, Surbhi Goel
We study online learning in the adversarial injection model introduced by [Goel et al. 2017], where a stream of labeled examples is predominantly drawn i.i.d.\ from an unknown dist…
cs.LG2026
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
Parsa Mirtaheri, Ezra Edelman, Samy Jelassi +2
Inference-time computation has emerged as a promising scaling axis for improving large language model reasoning. However, despite yielding impressive performance, the optimal alloc…