2 papers
math.OC2026
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
Curtis Fox, Aaron Mishkin, Sharan Vaswani +1
Iteration complexities for optimizing smooth functions with first-order algorithms are typically stated in terms of a global Lipschitz constant of the gradient, and near-optimal re…
cs.LG2025
Next-token prediction capacity: general upper bounds and a lower bound for transformers
Liam Madden, Curtis Fox, Christos Thrampoulidis
Given a sequence of tokens, such as words, the task of next-token prediction is to predict the next-token conditional probability distribution. Decoder-only transformers have becom…