57 citations · 72 across the 5 of their papers we have counts for
6 papers · 1 filter
TinyGSM: achieving >80% on GSM8k with small language models
Bingbin Liu, Sebastien Bubeck, Ronen Eldan +5
Small-scale models offer various computational advantages, and yet to which extent size is critical for problem-solving abilities remains an open question. Specifically for solving…
A law of robustness for two-layers neural networks
Sébastien Bubeck, Yuanzhi Li, Dheeraj Nagaraj
We initiate the study of the inherent tradeoffs between the size of a neural network and its robustness, as measured by its Lipschitz constant. We make a precise conjecture that, f…
Non-Stochastic Multi-Player Multi-Armed Bandits: Optimal Rate With Collision Information, Sublinear Without
Sébastien Bubeck, Yuanzhi Li, Yuval Peres +1
We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore…
Improved Path-length Regret Bounds for Bandits
Sébastien Bubeck, Yuanzhi Li, Haipeng Luo +1
We study adaptive regret bounds in terms of the variation of the losses (the so-called path-length bounds) for both multi-armed bandit and more generally linear bandit. We first sh…
Make the Minority Great Again: First-Order Regret Bound for Contextual Bandits
Zeyuan Allen-Zhu, Sébastien Bubeck, Yuanzhi Li
Regret bounds in online learning compare the player's performance to , the optimal performance in hindsight with a fixed strategy. Typically such bounds scale with the square…
Sparsity, variance and curvature in multi-armed bandits
Sébastien Bubeck, Michael B. Cohen, Yuanzhi Li
In (online) learning theory the concepts of sparsity, variance and curvature are well-understood and are routinely used to obtain refined regret and generalization bounds. In this…