198 citations
- Carnegie Mellon UniversityUS13 papers
- École PolytechniqueFR8 papers
- Australian National UniversityAU6 papers
- Inception Institute of Artificial IntelligenceAE6 papers
- Nanyang Technological UniversitySG6 papers
- Centre de Mathématiques Appliquées de l'École polytechniqueFR5 papers
- Centre National de la Recherche ScientifiqueFR5 papers
- Technical University of MunichDE5 papers
- Tohoku UniversityJP5 papers
- Hamad bin Khalifa UniversityQA4 papers
- Harbin Institute of TechnologyCN4 papers
- Harvard University PressUS4 papers
7 papers · 1 filter
Fast and Robust Likelihood-Guided Diffusion Posterior Sampling with Amortized Variational Inference
Léon Zheng, Thomas Hirtz, Yazid Janati +1
Zero-shot diffusion posterior sampling offers a flexible framework for inverse problems by accommodating arbitrary degradation operators at test time, but incurs high computational…
An Efficient Algorithm for Thresholding Monte Carlo Tree Search
Shoma Nameki, Atsuyoshi Nakamura, Junpei Komiyama +1
We introduce the Thresholding Monte Carlo Tree Search problem, in which, given a tree and a threshold , a player must answer whether the root node value of $\mathc…
Proximal Point Nash Learning from Human Feedback
Daniil Tiapkin, Daniele Calandriello, Denis Belomestny +5
Traditional Reinforcement Learning from Human Feedback (RLHF) often relies on reward models, frequently assuming preference structures like the Bradley--Terry model, which may not…
Scaffold with Stochastic Gradients: New Analysis with Linear Speed-Up
Paul Mangold, Alain Durmus, Aymeric Dieuleveut +1
This paper proposes a novel analysis for the Scaffold algorithm, a popular method for dealing with data heterogeneity in federated learning. While its convergence in deterministic…
Piecewise deterministic generative models
Andrea Bertazzi, Dario Shariatian, Umut Simsekli +2
We introduce a novel class of generative models based on piecewise deterministic Markov processes (PDMPs), a family of non-diffusive stochastic processes consisting of deterministi…
Model-free Posterior Sampling via Learning Rate Randomization
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +6
In this paper, we introduce Randomized Q-learning (RandQL), a novel randomized model-free algorithm for regret minimization in episodic Markov Decision Processes (MDPs). To the bes…