Showing cs.LGShow all
2 papers · 1 filter
cs.LG2024
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
Brian Cho, Dominik Meier, Kyra Gan +1
In multi-armed bandits, the tasks of reward maximization and pure exploration are often at odds with each other. The former focuses on exploiting arms with the highest means, while…
cs.LG2024
CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold Policies
Brian M Cho, Ana-Roxana Pop, Kyra Gan +4
When modifying existing policies in high-risk settings, it is often necessary to ensure with high certainty that the newly proposed policy improves upon a baseline, such as the sta…