11 citations · 21 across the 5 of their papers we have counts for
5 papers
gym-DSSAT: a crop model turned into a Reinforcement Learning environment
Romain Gautron, Emilio J. Padrón, Philippe Preux +3
Addressing a real world sequential decision problem with Reinforcement Learning (RL) usually starts with the use of a simulated environment that mimics real conditions. We present…
Indexed Minimum Empirical Divergence for Unimodal Bandits
Hassan Saber, Pierre Ménard, Odalric-Ambrym Maillard
We consider a multi-armed bandit problem specified by a set of one-dimensional family exponential distributions endowed with a unimodal structure. We introduce IMED-UB, a algorithm…
Random Shuffling and Resets for the Non-stationary Stochastic Bandit Problem
Robin Allesiardo, Raphaël Féraud, Odalric-Ambrym Maillard
We consider a non-stationary formulation of the stochastic multi-armed bandit where the rewards are no longer assumed to be identically distributed. For the best-arm identification…
Low-rank Bandits with Latent Mixtures
Aditya Gopalan, Odalric-Ambrym Maillard, Mohammadi Zaki
We study the task of maximizing rewards from recommending items (actions) to users sequentially interacting with a recommender system. Users are modeled as latent mixtures of C man…
Selecting Near-Optimal Approximate State Representations in Reinforcement Learning
Ronald Ortner, Odalric-Ambrym Maillard, Daniil Ryabko
We consider a reinforcement learning setting introduced in (Maillard et al., NIPS 2011) where the learner does not have explicit access to the states of the underlying Markov decis…