18 papers
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification
Alex Buna, Shirley Xiaoqi Liu, Patrick Rebeschini
In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic lo…
Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning
Hen Davidov, Nachshon Cohen, Oren Kalinsky +4
LLMs utilizing chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can mitigate this by withholding outputs unlikely to be…
Aggregation with Exponential Weights is Optimal in Expectation
Mikael Møller Høgsgaard, Patrick Rebeschini, Tobias Wegel
The aggregation with exponential weights (AEW) estimator is not fully understood in the basic setting of model selection aggregation with squared loss. In particular, whether it is…
Masked Language Flow Models
Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang +2
Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in th…
Self-Concordant Perturbations for Linear Bandits
Lucas Lévy, Jean-Lou Valeau, Arya Akhavan +1
We consider the adversarial linear bandits setting and present a unified algorithmic framework that bridges Follow-the-Regularized-Leader (FTRL) and Follow-the-Perturbed-Leader (FT…
Generalization in Nonlinear Least Squares via Learned Feature Geometry
Ayub Kharel, Ilja Kuzborskij, Patrick Rebeschini +1
We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-…