collaborators

8 papers

math.ST2026

A Note on k-NN Gating in RAG

Gérard Biau, Claire Boyer

We propose a statistical proxy framework for retrieval-augmented generation (RAG) that formalizes how language models balance internal predictions with retrieved evidence. We deriv…

stat.ML2026

Optimal Stopping in Latent Diffusion Models

Yu-Han Wu, Quentin Berthet, Gérard Biau +3

We identify and analyze a surprising phenomenon of Latent Diffusion Models (LDMs) where the final steps of the diffusion can degrade sample quality. In contrast to conventional arg…

cs.LG2026

Statistical Advantage of Softmax Attention: Insights from Single-Location Regression

O. Duranthon, P. Marion, C. Boyer +2

Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly underst…

stat.ML2025

Attention-based clustering

Rodrigo Maulen-Soto, Pierre Marion, Claire Boyer

Transformers have emerged as a powerful neural network architecture capable of tackling a wide range of learning tasks. In this work, we provide a theoretical analysis of their abi…

stat.ML2025

Fast kernel methods: Sobolev, physics-informed, and additive models

Nathan Doumèche, Francis Bach, Gérard Biau +1

Kernel methods are powerful tools in statistical learning, but their cubic complexity in the sample size n limits their use on large-scale datasets. In this work, we introduce a sc…

stat.ML2025

Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization

Yu-Han Wu, Pierre Marion, Gérard Biau +1

Denoising score matching plays a pivotal role in the performance of diffusion-based generative models. However, the empirical optimal score--the exact solution to the denoising sco…