5 papers
Minimax PAC Bounds for Learning in Exogenous Contextual MDPs
Corentin Pla, Hugo Richard, Marc Abeille +1
We study PAC learning in tabular discounted Markov decision processes with exogenous i.i.d. contexts, with discount factor , finite state space , action space $\mat…
When and why randomised exploration works (in linear bandits)
Marc Abeille, David Janz, Ciara Pike-Burke
We provide an approach for the analysis of randomised exploration algorithms like Thompson sampling that does not rely on forced optimism or posterior inflation. With this, we demo…
Variance-sensitive Thompson sampling for generalised linear bandits, revisited
Tom Perneczky, Marc Abeille, David Janz
We prove a variance-sensitive regret bound for Thompson sampling in stochastic generalised linear bandits. The argument assumes a warm-up, after which the regret is controlled thro…
On the Hardness of Reinforcement Learning with Transition Look-Ahead
Corentin Pla, Hugo Richard, Marc Abeille +2
We study reinforcement learning (RL) with transition look-ahead, where the agent may observe which states would be visited upon playing any sequence of actions before decidi…
Multi-Armed Bandits with Minimum Aggregated Revenue Constraints
Ahmed Ben Yahmed, Hafedh El Ferchichi, Marc Abeille +1
We examine a multi-armed bandit problem with contextual information, where the objective is to ensure that each arm receives a minimum aggregated reward across contexts while simul…