most citedRisk-Aversion in Multi-armed Bandits

93 citations · 204 across the 8 of their papers we have counts for

collaborators

8 papers

stat.ML2013

Toward Optimal Stratification for Stratified Monte-Carlo Integration

Alexandra Carpentier, Remi Munos

We consider the problem of adaptive stratified sampling for Monte Carlo integration of a noisy function, given a finite budget n of noisy evaluations to the function. We tackle in…

cs.LG201330 cited

Selecting the State-Representation in Reinforcement Learning

Odalric-Ambrym Maillard, Rémi Munos, Daniil Ryabko

The problem of selecting the right state-representation in a reinforcement learning problem is considered. Several models (functions mapping past observations to a finite set) of t…

cs.LG201393 cited

Risk-Aversion in Multi-armed Bandits

Amir Sani, Alessandro Lazaric, Rémi Munos

Stochastic multi-armed bandits solve the Exploration-Exploitation dilemma and ultimately maximize the expected reward. Nonetheless, in many practical problems, maximizing the expec…

stat.ML20127 cited

Adaptive Stratified Sampling for Monte-Carlo integration of Differentiable functions

Alexandra Carpentier, Rémi Munos

We consider the problem of adaptive stratified sampling for Monte Carlo integration of a differentiable function given a finite number of evaluations to the function. We construct…

cs.LG201232 cited

Regret Bounds for Restless Markov Bandits

Ronald Ortner, Daniil Ryabko, Peter Auer +1

We consider the restless Markov bandit problem, in which the state of each arm evolves according to a Markov process independently of the learner's actions. We suggest an algorithm…

cs.LG201242 cited

On the Sample Complexity of Reinforcement Learning with a Generative Model

Mohammad Gheshlaghi Azar, Remi Munos, Bert Kappen

We consider the problem of learning the optimal action-value function in the discounted-reward Markov decision processes (MDPs). We prove a new PAC bound on the sample-complexity o…