activity
20172023
most citedSparsity, variance and curvature in multi-armed bandits

57 citations · 71 across the 4 of their papers we have counts for

collaborators

10 papers

cs.LG2023

TinyGSM: achieving >80% on GSM8k with small language models

Bingbin Liu, Sebastien Bubeck, Ronen Eldan +5

Small-scale models offer various computational advantages, and yet to which extent size is critical for problem-solving abilities remains an open question. Specifically for solving…

cs.CL20232 cited

Positional Description Matters for Transformers Arithmetic

Ruoqi Shen, Sébastien Bubeck, Ronen Eldan +3

Transformers, central to the successes in modern Natural Language Processing, often falter on arithmetic tasks despite their vast capabilities --which paradoxically include remarka…

cs.LG2020

A law of robustness for two-layers neural networks

Sébastien Bubeck, Yuanzhi Li, Dheeraj Nagaraj

We initiate the study of the inherent tradeoffs between the size of a neural network and its robustness, as measured by its Lipschitz constant. We make a precise conjecture that, f…

math.OC2019

Complexity of Highly Parallel Non-Smooth Convex Optimization

Sébastien Bubeck, Qijia Jiang, Yin Tat Lee +2

A landmark result of non-smooth convex optimization is that gradient descent is an optimal algorithm whenever the number of computed gradients is smaller than the dimension . In…

cs.LG20195 cited

Non-Stochastic Multi-Player Multi-Armed Bandits: Optimal Rate With Collision Information, Sublinear Without

Sébastien Bubeck, Yuanzhi Li, Yuval Peres +1

We consider the non-stochastic version of the (cooperative) multi-player multi-armed bandit problem. The model assumes no communication at all between the players, and furthermore…

cs.LG20197 cited

Improved Path-length Regret Bounds for Bandits

Sébastien Bubeck, Yuanzhi Li, Haipeng Luo +1

We study adaptive regret bounds in terms of the variation of the losses (the so-called path-length bounds) for both multi-armed bandit and more generally linear bandit. We first sh…