12 citations · 26 across the 23 of their papers we have counts for
5 papers · 1 filter
Minimax Optimal Early-Stopped Gradient Descent for Gaussian Mixture Classification
Alex Buna, Shirley Xiaoqi Liu, Patrick Rebeschini
In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic lo…
Aggregation with Exponential Weights is Optimal in Expectation
Mikael Møller Høgsgaard, Patrick Rebeschini, Tobias Wegel
The aggregation with exponential weights (AEW) estimator is not fully understood in the basic setting of model selection aggregation with squared loss. In particular, whether it is…
Masked Language Flow Models
Iskander Azangulov, Kianoosh Ashouritaklimi, Leo Zhang +2
Masked Diffusion Models (MDMs) promise fast, parallel language generation, but their reverse transition factorises across token positions -- an approximation that breaks down in th…
Generalization in Nonlinear Least Squares via Learned Feature Geometry
Ayub Kharel, Ilja Kuzborskij, Patrick Rebeschini +1
We study the generalization of ridge-regularized nonlinear least-squares models via on-average algorithmic stability, deriving error bounds for local minimizers in terms of a data-…
On-Average Stability of Multipass Preconditioned SGD and Effective Dimension
Simon Vary, Tyler Farghly, Ilja Kuzborskij +1
We study trade-offs between the population risk curvature, geometry of the noise, and preconditioning on the generalisation ability of the multipass Preconditioned Stochastic Gradi…