5 papers
Bandit Simulation for Average Reward Inference
Samya Praharaj, Chih-Yu Chang, Koulik Khamaru +1
Multi-arm bandit algorithms are increasingly used in online platforms, clinical trials, and social science experiments, but valid statistical inference on their performance remains…
TRAP: Tail-aware Ranking Attack for World-Model Planning
Siyuan Duan, Ke Zhang, Xizhao Luo
World models enable long-horizon planning by internally generating and evaluating imagined trajectories, making them a promising foundation for generalist agents. However, this ima…
Contextual Thompson Sampling via Generation of Missing Data
Kelly W. Zhang, Tiffany Tianhui Cai, Hongseok Namkoong +1
We introduce a framework for Thompson sampling (TS) contextual bandit algorithms, in which the algorithm's ability to quantify uncertainty and make decisions depends on the quality…
Impatient Bandits: Optimizing for the Long-Term Without Delay
Kelly W. Zhang, Thomas Baldwin-McDonald, Kamil Ciosek +2
Increasingly, recommender systems are tasked with improving users' long-term satisfaction. In this context, we study a content exploration task, which we formalize as a bandit prob…
Active Exploration via Autoregressive Generation of Missing Data
Tiffany Tianhui Cai, Hongseok Namkoong, Daniel Russo +1
We pose uncertainty quantification and exploration in online decision-making as a problem of training and generation from an autoregressive sequence model, an area experiencing rap…