2 papers
stat.ML2026
The Greedy Advantage in Finite-Horizon Bandits
Kai Zhou, Michael Lingzhi Li, Kai Wang
Organizations increasingly rely on sequential experimentation to improve decision-making. While the multi-armed bandit literature has developed algorithms with strong asymptotic re…
cs.LG2026
Learning to Cover: Online Learning and Optimization with Irreversible Decisions
Alexandre Jacquillat, Michael Lingzhi Li
We define an online learning and optimization problem with discrete and irreversible decisions contributing toward a coverage target. In each period, a decision-maker selects facil…