collaborators

10 papers

math.ST2026

Adjusted Scores for Discrete Langevin Algorithms

Armand Gissler, Saeed Saremi, Francis Bach

Sampling from discrete distributions is a ubiquitous task in machine learning, recently revisited by the emergence of discrete diffusion models. While Langevin algorithms constitut…

math.OC2026

Two-stage stochastic algorithm for solving large-scale (non)-convex separable optimization problems under affine constraints

Benjamin Dubois-Taine, Laurent Pfeiffer, Nadia Oudjane +2

We consider nonsmooth optimization problems under affine constraints, where the objective consists of the average of the component functions of a large number of agents, and we…

cs.LG2025

A Convex Loss Function for Set Prediction with Optimal Trade-offs Between Size and Conditional Coverage

Francis Bach

We consider supervised learning problems in which set predictions provide explicit uncertainty estimates. Using Choquet integrals (a.k.a. Lov{á}sz extensions), we propose a convex…

stat.ML2025

Convergence of Shallow ReLU Networks on Weakly Interacting Data

Léo Dana, Francis Bach, Loucas Pillaud-Vivien

We analyse the convergence of one-hidden-layer ReLU networks trained by gradient flow on data points. Our main contribution leverages the high dimensionality of the ambient spa…

cs.LG2025

On the Effectiveness of the z-Transform Method in Quadratic Optimization

Francis Bach

The z-transform of a sequence is a classical tool used within signal processing, control theory, computer science, and electrical engineering. It allows for studying sequences from…

cs.LG2025

Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law

Frederik Kunstner, Francis Bach

Recent works have highlighted optimization difficulties faced by gradient descent in training the first and last layers of transformer-based language models, which are overcome by…