Showing stat.MLShow all
2 papers · 1 filter
stat.ML2026
Online Statistical Inference of Constant Sample-averaged Q-Learning
Saunak Kumar Panda, Tong Li, Ruiqi Liu +1
Reinforcement learning algorithms have been widely used for decision-making tasks in various domains. However, the performance of these algorithms can be impacted by high variance…
stat.ML2026
RIE-Greedy: Regularization-Induced Exploration for Contextual Bandits
Tong Li, Thiago de Queiroz Casanova, Eric M. Schwartz +3
Real-world contextual bandit problems with complex reward models are often tackled with iteratively trained models, such as boosting trees. However, it is difficult to directly app…