Publications (4)
Flatland: The Adventures of Gradient Descent with Large Step Sizes
Leonardo Galli, Curtis Fox, Wiebke Bartolomaeus +2
The training of neural networks often entails objective functions that are not globally -smooth. For these functions, it is both theoretically and practically difficult to reply…
Next-token prediction capacity: general upper bounds and a lower bound for transformers
Liam Madden, Curtis Fox, Christos Thrampoulidis
Given a sequence of tokens, such as words, the task of next-token prediction is to predict the next-token conditional probability distribution. Decoder-only transformers have becom…
Glocal Smoothness: Line search and adaptive step sizes can help in theory too!
Curtis Fox, Aaron Mishkin, Sharan Vaswani +1
Iteration complexities for optimizing smooth functions with first-order algorithms are typically stated in terms of a global Lipschitz constant of the gradient, and near-optimal re…
DRBench: A Realistic Benchmark for Enterprise Deep Research
Amirhossein Abaskohi, Tianyi Chen, Miguel Muñoz-Mármol +11
We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions…