2 papers
stat.ML2023
Fast Rates for Maximum Entropy Exploration
Daniil Tiapkin, Denis Belomestny, Daniele Calandriello +7
We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maxim…
stat.ML2023
When Combinatorial Thompson Sampling meets Approximation Regret
Pierre Perrault
We study the Combinatorial Thompson Sampling policy (CTS) for combinatorial multi-armed bandit problems (CMAB), within an approximation regret setting. Although CTS has attracted a…