papers

Publications (37)

cs.LG2020

Learning with Differentiable Perturbed Optimizers

Quentin Berthet, Mathieu Blondel, Olivier Teboul +3

Machine learning pipelines often rely on optimization procedures to make discrete decisions (e.g., sorting, picking closest neighbors, or shortest paths). Although these discrete d…

stat.ML2026

Optimal Stopping in Latent Diffusion Models

Yu-Han Wu, Quentin Berthet, Gérard Biau +3

We identify and analyze a surprising phenomenon of Latent Diffusion Models (LDMs) where the final steps of the diffusion can degrade sample quality. In contrast to conventional arg…

cs.LG2022

Efficient and Modular Implicit Differentiation

Mathieu Blondel, Quentin Berthet, Marco Cuturi +5

Automatic differentiation (autodiff) has revolutionized machine learning. It allows to express complex computations by composing elementary ones in creative ways and removes the bu…

cs.LG2023

Differentiable Clustering with Perturbed Spanning Forests

Lawrence Stewart, Francis S Bach, Felipe Llinares López +1

We introduce a differentiable clustering method based on stochastic perturbations of minimum-weight spanning forests. This allows us to include clustering in end-to-end trainable p…

math.ST2017

Exact recovery in the Ising blockmodel

Quentin Berthet, Philippe Rigollet, Piyush Srivastava

We consider the problem associated to recovering the block structure of an Ising model given independent observations on the binary hypercube. This new model, called the Ising bloc…

cs.LG2025

Control Variate Score Matching for Diffusion Models

Khaled Kahouli, Romuald Elie, Klaus-Robert Müller +3

Diffusion models offer a robust framework for sampling from unnormalized probability densities, which requires accurately estimating the score of the noise-perturbed target distrib…

math.ST2016

Statistical and computational trade-offs in estimation of sparse principal components

Tengyao Wang, Quentin Berthet, Richard J. Samworth

In recent years, sparse principal component analysis has emerged as an extremely popular dimension reduction technique for high-dimensional data. The theoretical challenge, in the…

math.ST2018

Statistical Windows in Testing for the Initial Distribution of a Reversible Markov Chain

Quentin Berthet, Varun Kanade

We study the problem of hypothesis testing between two discrete distributions, where we only have access to samples after the action of a known reversible Markov chain, playing the…

math.OC2024

Fast Stochastic Composite Minimization and an Accelerated Frank-Wolfe Algorithm under Parallelization

Benjamin Dubois-Taine, Francis Bach, Quentin Berthet +1

We consider the problem of minimizing the sum of two convex functions. One of those functions has Lipschitz-continuous gradients, and can be accessed via stochastic oracles, wherea…

stat.ML2020

Fast Differentiable Sorting and Ranking

Mathieu Blondel, Olivier Teboul, Quentin Berthet +1

The sorting operation is one of the most commonly used building blocks in computer programming. In machine learning, it is often used for robust statistics. However, seen as a func…

cs.LG2026

Diffusion Fine-tuning with Rewarded Moment Matching Distillation

Alexis Jacq, Guillaume Couairon, Valentin De Bortoli +3

Distillation and Reinforcement Learning (RL) fine-tuning are the primary pillars of diffusion post-training. While traditionally studied in isolation, the interaction between these…

stat.ME2020

Noisy Adaptive Group Testing using Bayesian Sequential Experimental Design

Marco Cuturi, Olivier Teboul, Quentin Berthet +2

When the infection prevalence of a disease is low, Dorfman showed 80 years ago that testing groups of people can prove more efficient than testing people individually. Our goal in…

cs.LG2023

Regression as Classification: Influence of Task Formulation on Neural Network Features

Lawrence Stewart, Francis Bach, Quentin Berthet +1

Neural networks can be trained to solve regression problems by using gradient-based methods to minimize the square loss. However, practitioners often prefer to reformulate regressi…

cs.CL2026

DiffusionGemma Technical Report

DiffusionGemma Team, Adrien Ali Taïga, James Assiene +41

We introduce DiffusionGemma, an experimental open-weight language model that uses discrete diffusion to generate text at exceptionally high speed. Rather than decoding one token at…

cs.SD2021

Self-Supervised Learning of Audio Representations from Permutations with Differentiable Ranking

Andrew N Carr, Quentin Berthet, Mathieu Blondel +2

Self-supervised pre-training using so-called "pretext" tasks has recently shown impressive performance across a wide range of modalities. In this work, we advance self-supervised l…

cs.LG2018

Unsupervised Alignment of Embeddings with Wasserstein Procrustes

Edouard Grave, Armand Joulin, Quentin Berthet

We consider the task of aligning two sets of points in high dimension, which has many applications in natural language processing and computer vision. As an example, it was recentl…

cs.LG2017

Fast Rates for Bandit Optimization with Upper-Confidence Frank-Wolfe

Quentin Berthet, Vianney Perchet

We consider the problem of bandit optimization, inspired by stochastic optimization and online learning problems with bandit feedback. In this problem, the objective is to minimize…

stat.AP2014

Resource Allocation for Statistical Estimation

Quentin Berthet, Venkat Chandrasekaran

Statistical estimation in many contemporary settings involves the acquisition, analysis, and aggregation of datasets from multiple sources, which can have significant differences i…

math.ST2013

Computational Lower Bounds for Sparse PCA

Quentin Berthet, Philippe Rigollet

In the context of sparse principal component detection, we bring evidence towards the existence of a statistical price to pay for computational efficiency. We measure the performan…

stat.ML2019

Regularized Contextual Bandits

Xavier Fontaine, Quentin Berthet, Vianney Perchet

We consider the stochastic contextual bandit problem with additional regularization. The motivation comes from problems where the policy of the agent must be close to some baseline…

cs.LG2026

MIND: Monge Inception Distance for Generative Models Evaluation

Quentin Berthet, Yu-Han Wu, Clement Crepy +3

We propose the Monge Inception Distance (MIND), a metric for evaluating generative models that addresses key limitations of the widely adopted Fréchet Inception Distance (FID). Th…

math.ST2015

Optimal Testing for Planted Satisfiability Problems

Quentin Berthet

We study the problem of detecting planted solutions in a random satisfiability formula. Adopting the formalism of hypothesis testing in statistical analysis, we describe the minima…

cs.MA2025

Soft Condorcet Optimization for Ranking of General Agents

Marc Lanctot, Kate Larson, Michael Kaisers +7

Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a po…

cs.LG2025

Implicit Diffusion: Efficient Optimization through Stochastic Sampling

Pierre Marion, Anna Korba, Peter Bartlett +6

We present a new algorithm to optimize distributions defined implicitly by parameterized stochastic diffusions. Doing so allows us to modify the outcome distribution of sampling pr…

cs.LG2024

Decoding-time Realignment of Language Models

Tianlin Liu, Shangmin Guo, Leonardo Bianco +7

Aligning language models with human preferences is crucial for reducing errors and biases in these models. Alignment techniques, such as reinforcement learning from human feedback…

cs.NE2022

A unified software/hardware scalable architecture for brain-inspired computing based on self-organizing neural models

Artem R. Muliukov, Laurent Rodriguez, Benoit Miramond +4

The field of artificial intelligence has significantly advanced over the past decades, inspired by discoveries from the fields of biology and neuroscience. The idea of this work is…

cs.LG2026

Kastor: An efficient fine-tuning strategy for generative emulation of PDE simulations

Guillaume Couairon, Alexis Jacq, Yu-Han Wu +4

Machine learning offers a promising avenue to accelerate physical simulations by replacing computationally expensive traditional Partial Differential Equation (PDE) solvers with fa…

cs.LG2023

Mirror Sinkhorn: Fast Online Optimization on Transport Polytopes

Marin Ballu, Quentin Berthet

Optimal transport is an important tool in machine learning, allowing to capture geometric properties of the data through a linear program on transport polytopes. We present a singl…

math.ST2019

Detection of Planted Solutions for Flat Satisfiability Problems

Quentin Berthet, Jordan S. Ellenberg

We study the detection problem of finding planted solutions in random instances of flat satisfiability problems, a generalization of boolean satisfiability formulas. We describe th…

cs.LG2020

Stochastic Optimization for Regularized Wasserstein Estimators

Marin Ballu, Quentin Berthet, Francis Bach

Optimal transport is a foundational problem in optimization, that allows to compare probability distributions while taking into account geometric aspects. Its optimal objective val…

cs.LG2016

Average-case Hardness of RIP Certification

Tengyao Wang, Quentin Berthet, Yaniv Plan

The restricted isometry property (RIP) for design matrices gives guarantees for optimal recovery in sparse linear models. It is of high interest in compressed sensing and statistic…

stat.ML2026

Generalization Properties of Score-matching Diffusion Models for Intrinsically Low-dimensional Data

Saptarshi Chakraborty, Quentin Berthet, Peter L. Bartlett

Despite the remarkable empirical success of score-based diffusion models, their statistical guarantees remain underdeveloped. Existing analyses often provide pessimistic convergenc…

stat.ML2025

Building Bridges between Regression, Clustering, and Classification

Lawrence Stewart, Francis Bach, Quentin Berthet

Regression, the task of predicting a continuous scalar target y based on some features x is one of the most fundamental tasks in machine learning and statistics. It has been observ…

math.ST2020

Minimax estimation of smooth densities in Wasserstein distance

Jonathan Niles-Weed, Quentin Berthet

We study nonparametric density estimation problems where error is measured in the Wasserstein distance, a metric on probability distributions popular in many areas of statistics an…

math.ST2018

Optimal link prediction with matrix logistic regression

Nicolai Baldin, Quentin Berthet

We consider the problem of link prediction, based on partial observation of a large network, and on side information associated to its vertices. The generative model is formulated…

cs.AR2025

hls4ml: A Flexible, Open-Source Platform for Deep Learning Acceleration on Reconfigurable Hardware

Jan-Frederik Schulte, Benjamin Ramhorst, Chang Sun +50

We present hls4ml, a free and open-source platform that translates machine learning (ML) models from modern deep learning frameworks into high-level synthesis (HLS) code that can b…

math.ST2013

Optimal detection of sparse principal components in high dimension

Quentin Berthet, Philippe Rigollet

We perform a finite sample analysis of the detection levels for sparse principal components of a high-dimensional covariance matrix. Our minimax optimal test is based on a sparse e…