407 citations
- University of AmsterdamNL61 papers
- Eindhoven University of TechnologyNL60 papers
- Vrije Universiteit AmsterdamNL17 papers
- Radboud University NijmegenNL15 papers
- Centre National de la Recherche ScientifiqueFR13 papers
- Leiden UniversityNL13 papers
- College of Western IdahoUS12 papers
- University of WaterlooCA10 papers
- Delft University of TechnologyNL8 papers
- University of OxfordGB8 papers
- University of California, BerkeleyUS7 papers
- University of CambridgeGB7 papers
15 papers · 1 filter
Regret Minimization in Heavy-Tailed Bandits
Shubhada Agrawal, Sandeep Juneja, Wouter M. Koolen
We revisit the classic regret-minimization problem in the stochastic multi-armed bandit setting when the arm-distributions are allowed to be heavy-tailed. Regret minimization has b…
CDT: Cascading Decision Trees for Explainable Reinforcement Learning
Zihan Ding, Pablo Hernandez-Leal, Gavin Weiguang Ding +2
Deep Reinforcement Learning (DRL) has recently achieved significant advances in various domains. However, explaining the policy of RL agents still remains an open problem due to se…
Probabilistic Super-Resolution of Solar Magnetograms: Generating Many Explanations and Measuring Uncertainties
Xavier Gitiaux, Shane A. Maloney, Anna Jungbluth +7
Machine learning techniques have been successfully applied to super-resolution tasks on natural images where visually pleasing results are sufficient. However in many scientific do…
Fixed-Confidence Guarantees for Bayesian Best-Arm Identification
Xuedong Shang, Rianne de Heide, Emilie Kaufmann +2
We investigate and provide new insights on the sampling rule called Top-Two Thompson Sampling (TTTS). In particular, we justify its use for fixed-confidence best-arm identification…
Approximate Dynamic Programming with Neural Networks in Linear Discrete Action Spaces
Wouter van Heeswijk, Han La Poutré
Real-world problems of operations research are typically high-dimensional and combinatorial. Linear programs are generally used to formulate and efficiently solve these large decis…
Lipschitz Adaptivity with Multiple Learning Rates in Online Learning
Zakaria Mhammedi, Wouter M. Koolen, Tim van Erven
We aim to design adaptive online learning algorithms that take advantage of any special structure that might be present in the learning task at hand, with as little manual tuning b…