129 citations · 326 across the 9 of their papers we have counts for
8 papers · 1 filter
Bandits attack function optimization
Philippe Preux, Rémi Munos, Michal Valko
We consider function optimization as a sequential decision making problem under budget constraint. This constraint limits the number of objective function evaluations allowed durin…
Efficient learning by implicit exploration in bandit problems with side observations
Tomas Kocak, Gergely Neu, Michal Valko +1
We consider online learning problems under a partial observability model capturing situations where the information conveyed to the learner is between full information and bandit f…
Stochastic simultaneous optimistic optimization
Michal Valko, Alexandra Carpentier, Rémi Munos
We study the problem of global maximization of a function f given a finite number of evaluations perturbed by noise. We consider a very weak assumption on the function, namely that…
Planning in entropy-regularized Markov decision processes and games
Jean-Bastien Grill, Omar Darwiche Domingues, Pierre Ménard +2
We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model…
Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
Jean-Bastien Grill, Michal Valko, Rémi Munos
You are a robot and you live in a Markov decision process (MDP) with a finite or an infinite number of transitions from state-action to next states. You got brains and so you plan…
Spectral Thompson sampling
Tomas Kocak, Michal Valko, Remi Munos +1
Thompson Sampling (TS) has attracted a lot of interest due to its good empirical performance, in particular in the computational advertising. Though successful, the tools for its p…