6 papers
Adaptive Calibration in Non-Stationary Environments
Junyan Liu, Haipeng Luo, Lillian J. Ratliff
Making calibrated online predictions is a central challenge in modern AI systems. Much of the existing literature focuses on fully adversarial environments where outcomes may be ar…
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
Soumita Hait, Ping Li, Haipeng Luo +1
Last-iterate convergence of learning dynamics in games has attracted significant recent attention. In two-player zero-sum games with bandit feedback, where only the loss of the sel…
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
Mengxiao Zhang, Yuheng Zhang, Haipeng Luo +1
In this paper, we study Interaction-Grounded Learning (IGL) [Xie et al., 2021], a paradigm designed for realistic scenarios where the learner receives indirect feedback generated b…
Comparator-Adaptive -Regret: Improved Bounds, Simpler Algorithms, and Applications to Games
Soumita Hait, Ping Li, Haipeng Luo +1
In the classic expert problem, -regret measures the gap between the learner's total loss and that achieved by applying the best action transformation . A recent work…
Alternating Regret for Online Convex Optimization
Soumita Hait, Ping Li, Haipeng Luo +1
Motivated by alternating learning dynamics in two-player games, a recent work by Cevher et al.(2024) shows that alternating regret is possible for any -round adver…
Contextual Linear Bandits with Delay as Payoff
Mengxiao Zhang, Yingfei Wang, Haipeng Luo
A recent work by Schlisselberg et al. (2024) studies a delay-as-payoff model for stochastic multi-armed bandits, where the payoff (either loss or reward) is delayed for a period th…