3 papers
stat.ML2026
Optimal Stopping in Latent Diffusion Models
Yu-Han Wu, Quentin Berthet, Gérard Biau +3
We identify and analyze a surprising phenomenon of Latent Diffusion Models (LDMs) where the final steps of the diffusion can degrade sample quality. In contrast to conventional arg…
cs.LG2026
Statistical Advantage of Softmax Attention: Insights from Single-Location Regression
O. Duranthon, P. Marion, C. Boyer +2
Large language models rely on attention mechanisms with a softmax activation. Yet the dominance of softmax over alternatives (e.g., component-wise or linear) remains poorly underst…
stat.ML2025
Attention-based clustering
Rodrigo Maulen-Soto, Pierre Marion, Claire Boyer
Transformers have emerged as a powerful neural network architecture capable of tackling a wide range of learning tasks. In this work, we provide a theoretical analysis of their abi…