activity
20242026
collaborators

10 papers

cs.LG2026

Near-Optimal Stochastic Linear Bandits with Delay

Ofir Schlisselberg, Mengxiao Zhang, Yishay Mansour

We study stochastic linear bandits with delayed feedback under several delay models and establish near-optimal regret guarantees. Our results identify when delayed linear bandits e…

cs.LG2026

Near-Optimal Last-Iterate Convergence for Zero-Sum Games with Bandit Feedback and Opponent Actions

Soumita Hait, Ping Li, Haipeng Luo +1

Last-iterate convergence of learning dynamics in games has attracted significant recent attention. In two-player zero-sum games with bandit feedback, where only the loss of the sel…

cs.LG2026

Interaction-Grounded Learning for Contextual Markov Decision Processes with Personalized Feedback

Mengxiao Zhang, Yuheng Zhang, Haipeng Luo +1

In this paper, we study Interaction-Grounded Learning (IGL) [Xie et al., 2021], a paradigm designed for realistic scenarios where the learner receives indirect feedback generated b…

cs.LG2025

Comparator-Adaptive -Regret: Improved Bounds, Simpler Algorithms, and Applications to Games

Soumita Hait, Ping Li, Haipeng Luo +1

In the classic expert problem, -regret measures the gap between the learner's total loss and that achieved by applying the best action transformation . A recent work…

cs.CV2025

Defective Convolutional Networks

Tiange Luo, Tianle Cai, Mengxiao Zhang +3

Robustness of convolutional neural networks (CNNs) has gained in importance on account of adversarial examples, i.e., inputs added as well-designed perturbations that are impercept…

cs.LG2025

Alternating Regret for Online Convex Optimization

Soumita Hait, Ping Li, Haipeng Luo +1

Motivated by alternating learning dynamics in two-player games, a recent work by Cevher et al.(2024) shows that alternating regret is possible for any -round adver…