10 papers
Adjusted Scores for Discrete Langevin Algorithms
Armand Gissler, Saeed Saremi, Francis Bach
Sampling from discrete distributions is a ubiquitous task in machine learning, recently revisited by the emergence of discrete diffusion models. While Langevin algorithms constitut…
Two-stage stochastic algorithm for solving large-scale (non)-convex separable optimization problems under affine constraints
Benjamin Dubois-Taine, Laurent Pfeiffer, Nadia Oudjane +2
We consider nonsmooth optimization problems under affine constraints, where the objective consists of the average of the component functions of a large number of agents, and we…
A Convex Loss Function for Set Prediction with Optimal Trade-offs Between Size and Conditional Coverage
Francis Bach
We consider supervised learning problems in which set predictions provide explicit uncertainty estimates. Using Choquet integrals (a.k.a. Lov{á}sz extensions), we propose a convex…
Convergence of Shallow ReLU Networks on Weakly Interacting Data
Léo Dana, Francis Bach, Loucas Pillaud-Vivien
We analyse the convergence of one-hidden-layer ReLU networks trained by gradient flow on data points. Our main contribution leverages the high dimensionality of the ambient spa…
On the Effectiveness of the z-Transform Method in Quadratic Optimization
Francis Bach
The z-transform of a sequence is a classical tool used within signal processing, control theory, computer science, and electrical engineering. It allows for studying sequences from…
Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
Frederik Kunstner, Francis Bach
Recent works have highlighted optimization difficulties faced by gradient descent in training the first and last layers of transformer-based language models, which are overcome by…