10 papers
Near-Optimal Stochastic Linear Bandits with Delay
Ofir Schlisselberg, Mengxiao Zhang, Yishay Mansour
We study stochastic linear bandits with delayed feedback under several delay models and establish near-optimal regret guarantees. Our results identify when delayed linear bandits e…
Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions
Soumita Hait, Ping Li, Haipeng Luo +1
Last-iterate convergence of learning dynamics in games has attracted significant recent attention. In two-player zero-sum games with bandit feedback, where only the loss of the sel…
Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback
Mengxiao Zhang, Yuheng Zhang, Haipeng Luo +1
In this paper, we study Interaction-Grounded Learning (IGL) [Xie et al., 2021], a paradigm designed for realistic scenarios where the learner receives indirect feedback generated b…
Comparator-Adaptive -Regret: Improved Bounds, Simpler Algorithms, and Applications to Games
Soumita Hait, Ping Li, Haipeng Luo +1
In the classic expert problem, -regret measures the gap between the learner's total loss and that achieved by applying the best action transformation . A recent work…
Defective Convolutional Networks
Tiange Luo, Tianle Cai, Mengxiao Zhang +3
Robustness of convolutional neural networks (CNNs) has gained in importance on account of adversarial examples, i.e., inputs added as well-designed perturbations that are impercept…
Alternating Regret for Online Convex Optimization
Soumita Hait, Ping Li, Haipeng Luo +1
Motivated by alternating learning dynamics in two-player games, a recent work by Cevher et al.(2024) shows that alternating regret is possible for any -round adver…