3 papers
stat.ML2025
SNPL: Simultaneous Policy Learning and Evaluation for Safe Multi-Objective Policy Improvement
Brian Cho, Ana-Roxana Pop, Ariel Evnine +1
To design effective digital interventions, experimenters face the challenge of learning decision policies that balance multiple objectives using offline data. Often, they aim to de…
cs.LG2024
Reward Maximization for Pure Exploration: Minimax Optimal Good Arm Identification for Nonparametric Multi-Armed Bandits
Brian Cho, Dominik Meier, Kyra Gan +1
In multi-armed bandits, the tasks of reward maximization and pure exploration are often at odds with each other. The former focuses on exploiting arms with the highest means, while…
cs.LG2024
CSPI-MT: Calibrated Safe Policy Improvement with Multiple Testing for Threshold Policies
Brian M Cho, Ana-Roxana Pop, Kyra Gan +4
When modifying existing policies in high-risk settings, it is often necessary to ensure with high certainty that the newly proposed policy improves upon a baseline, such as the sta…