combinatorial semi-bandits 1contextual bandits 1function approximation 1online learning 1regret analysis 1
From the 1 of 3 linked papers with an AI index.
3 papers
cs.LG2026
Optimal and Efficient Contextual Combinatorial Semi-bandits with General Function Approximation
Hao Qin, Chicheng Zhang
The paper introduces SquareCB.Comb, an efficient algorithm for contextual combinatorial semi‑bandits with general reward function approximation, achieving minimax optimal regret wi…
cs.LG2026
Taming the Monster Every Context: Complexity Measure and Unified Framework for Offline-Oracle Efficient Contextual Bandits
Hao Qin, Chicheng Zhang
We propose an algorithmic framework, Offline Estimation to Decisions (OE2D), that efficiently reduces contextual bandit learning with general reward function approximation to offli…
cs.LG2025
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
Hao Qin, Kwang-Sung Jun, Chicheng Zhang
We study the problem of -armed bandits with reward distributions belonging to a one-parameter exponential distribution family. In the literature, several criteria have been prop…