output
20022026
most citedA regression-based Monte Carlo method to solve backward stochastic differential equations

425 citations

Showing 2025 · stat.MLShow all

6 papers · 2 filters

stat.ML2025

Proximal Point Nash Learning from Human Feedback

Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5

Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…

stat.ML2025

Improving the evaluation of samplers on multi-modal targets

Louis Grenioux, Maxence Noble, Marylou Gabrié

Addressing multi-modality constitutes one of the major challenges of sampling. In this reflection paper, we advocate for a more systematic evaluation of samplers towards two source…

stat.ML2025

Personalized Convolutional Dictionary Learning of Physiological Time Series

Axel Roques, Samuel Gruffaz, Kyurae Kim +2

Human physiological signals tend to exhibit both global and local structures: the former are shared across a population, while the latter reflect inter-individual variability. For…

stat.ML2025

Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up

Paul Mangold, Alain Durmus, Aymeric Dieuleveut +1

This paper proposes a novel analysis for the Scaffold algorithm, a popular method for dealing with data heterogeneity in federated learning. While its convergence in deterministic…

stat.ML2025

Bit-Level Discrete Diffusion with Markov Probabilistic Models: An Improved Framework with Sharp Convergence Bounds under Minimal Assumptions

Le-Tuyet-Nhi Pham, Dario Shariatian, Antonio Ocello +2

This paper introduces Discrete Markov Probabilistic Models (DMPMs), a novel discrete diffusion algorithm for discrete data generation. The algorithm operates in discrete bit space,…

stat.ML2025

Beyond Log-Concavity and Score Regularity: Improved Convergence Bounds for Score-Based Generative Models in W2-distance

Marta Gentiloni-Silveri, Antonio Ocello

Score-based Generative Models (SGMs) aim to sample from a target distribution by learning score functions using samples perturbed by Gaussian noise. Existing convergence bounds for…