4 papers
Distilling LLM Feedback for Lean Theorem Proving
Gaetan Narozniak, Gérard Biau, Rémi Munos +2
Post-training for reasoning models typically combines supervised fine-tuning with reinforcement learning from verifiable rewards, most commonly with GRPO. However, this algorithm s…
Taking a Big Step: Large Learning Rates in Denoising Score Matching Prevent Memorization
Yu-Han Wu, Pierre Marion, Gérard Biau +1
Denoising score matching plays a pivotal role in the performance of diffusion-based generative models. However, the empirical optimal score--the exact solution to the denoising sco…
Scaling ResNets in the Large-depth Regime
Pierre Marion, Adeline Fermanian, Gérard Biau +1
Deep ResNets are recognized for achieving state-of-the-art results in complex machine learning tasks. However, the remarkable performance of these architectures relies on a trainin…
Attention layers provably solve single-location regression
Pierre Marion, Raphaël Berthier, Gérard Biau +1
Attention-based models, such as Transformer, excel across various tasks but lack a comprehensive theoretical understanding, especially regarding token-wise sparsity and internal li…