Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Statistical Inference for Misspecified Contextual Bandits
Yongyi Guo, Ziping Xu
Contextual bandit algorithms have transformed modern experimentation by enabling real-time adaptation for personalized treatment. Yet these advantages create challenges for statist…
stat.ML2024
The Fallacy of Minimizing Cumulative Regret in the Sequential Task Setting
Ziping Xu, Kelly W. Zhang, Susan A. Murphy
Online Reinforcement Learning (RL) is typically framed as the process of minimizing cumulative regret (CR) through interactions with an unknown environment. However, real-world RL…