Showing 2025 · cs.LGShow all
3 papers · 2 filters
cs.LG2025
Comparator-Adaptive -Regret: Improved Bounds, Simpler Algorithms, and Applications to Games
Soumita Hait, Ping Li, Haipeng Luo +1
In the classic expert problem, -regret measures the gap between the learner's total loss and that achieved by applying the best action transformation . A recent work by…
cs.LG2025
Contextual Linear Bandits with Delay as Payoff
Mengxiao Zhang, Yingfei Wang, Haipeng Luo
A recent work by Schlisselberg et al. (2024) studies a delay-as-payoff model for stochastic multi-armed bandits, where the payoff (either loss or reward) is delayed for a period th…
cs.LG2025
Alternating Regret for Online Convex Optimization
Soumita Hait, Ping Li, Haipeng Luo +1
Motivated by alternating learning dynamics in two-player games, a recent work by Cevher et al.(2024) shows that alternating regret is possible for any -round adver…