activity
20242026
collaborators

18 papers

stat.ML2026

Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification

Alex Buna, Shirley Xiaoqi Liu, Patrick Rebeschini

In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic lo…

cs.LG2026

Knowing When to Quit: A Principled Framework for Dynamic Abstention in LLM Reasoning

Hen Davidov, Nachshon Cohen, Oren Kalinsky +4

LLMs utilizing chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can mitigate this by withholding outputs unlikely to be…

math.ST2026

Aggregation with Exponential Weights is Optimal in Expectation

Mikael Møller Høgsgaard, Patrick Rebeschini, Tobias Wegel

The aggregation with exponential weights (AEW) estimator is not fully understood in the basic setting of model selection aggregation with squared loss. In particular, whether it is…

cs.CL2026

Masked Language Flow Models

Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang +2

Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in th…

stat.ML2026

Self-Concordant Perturbations for Linear Bandits

Lucas Lévy, Jean-Lou Valeau, Arya Akhavan +1

We consider the adversarial linear bandits setting and present a unified algorithmic framework that bridges Follow-the-Regularized-Leader (FTRL) and Follow-the-Perturbed-Leader (FT…

stat.ML2026

Generalization in Nonlinear Least Squares via Learned Feature Geometry

Ayub Kharel, Ilja Kuzborskij, Patrick Rebeschini +1

We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-…