84 citations
- Laboratoire de Probabilités et Modèles AléatoiresFR194 papers
- Sorbonne UniversitéFR76 papers
- Université Paris CitéFR59 papers
- Centre National de la Recherche ScientifiqueFR52 papers
- Sorbonne Paris CitéFR35 papers
- Institut Universitaire de FranceFR20 papers
- Laboratoire de Mathématiques Blaise PascalFR20 papers
- Laboratoire de Mathématiques d'OrsayFR20 papers
- École PolytechniqueFR17 papers
- Centre de Mathématiques Appliquées de l'École polytechniqueFR15 papers
- Université Paris-SaclayFR14 papers
- Département de mathématiques et applicationsFR12 papers
5 papers · 2 filters
Efficient Model-Based Concave Utility Reinforcement Learning through Greedy Mirror Descent
Bianca Marin Moreno, Margaux Brégère, Pierre Gaillard +1
Many machine learning tasks can be solved by minimizing a convex function of an occupancy measure over the policies that generate them. These include reinforcement learning, imitat…
Non-linear non-zero-sum Dynkin games with Bermudan strategies
Miryana Grigorova, Marie-Claire Quenez, Yuan Peng
In this paper, we study a non-zero-sum game with two players, where each of the players plays what we call Bermudan strategies and optimizes a general non-linear assessment functio…
Generative modeling for time series via Schr{ö}dinger bridge
Mohamed Hamdouche, Pierre Henry-Labordere, Huyên Pham
We propose a novel generative model for time series based on Schr{ö}dinger bridge (SB) approach. This consists in the entropic interpolation via optimal transport between a referen…
Discrete-time Mean-Field Stochastic Control with Partial Observations
Jeremy Chichportich, Idris Kharroubi
We study the optimal control of discrete time mean filed dynamical systems under partial observations. We express the global law of the filtered process as a controlled system with…
Non asymptotic analysis of Adaptive stochastic gradient algorithms and applications
Antoine Godichon-Baggioni, Pierre Tarrago
In stochastic optimization, a common tool to deal sequentially with large sample is to consider the well-known stochastic gradient algorithm. Nevertheless, since the stepsequence i…