Showing stat.MLShow all
3 papers · 1 filter
stat.ML2025
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
Yu-Han Wu, Pierre Marion, Gérard Biau +1
Denoising score matching plays a pivotal role in the performance of diffusion-based generative models. However, the empirical optimal score--the exact solution to the denoising sco…
stat.ML2025
Attention layers provably solve single-location regression
Pierre Marion, Raphaël Berthier, Gérard Biau +1
Attention-based models, such as Transformer, excel across various tasks but lack a comprehensive theoretical understanding, especially regarding token-wise sparsity and internal li…
stat.ML2024
Implicit regularization of deep residual networks towards neural ODEs
Pierre Marion, Yu-Han Wu, Michael E. Sander +1
Residual neural networks are state-of-the-art deep learning models. Their continuous-depth analog, neural ordinary differential equations (ODEs), are also widely used. Despite thei…