From the 1 of 4 linked papers with an AI index.
4 papers
Optimal and Efficient Contextual Combinatorial Semi-bandits with General Function Approximation
Hao Qin, Chicheng Zhang
The paper introduces SquareCB.Comb, an efficient algorithm for contextual combinatorial semi‑bandits with general reward function approximation, achieving minimax optimal regret wi…
Taming the Monster Every Context: Complexity Measure and Unified Framework for Offline-Oracle Efficient Contextual Bandits
Hao Qin, Chicheng Zhang
We propose an algorithmic framework, Offline Estimation to Decisions (OE2D), that efficiently reduces contextual bandit learning with general reward function approximation to offli…
Physics-Informed Parametric Bandits for Beam Alignment in mmWave Communications
Hao Qin, Thang Duong, Ming F. Li +1
In millimeter wave (mmWave) communications, beam alignment and tracking are crucial to combat the significant path loss. As scanning the entire directional space is inefficient, de…
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
Hao Qin, Kwang-Sung Jun, Chicheng Zhang
We study the problem of -armed bandits with reward distributions belonging to a one-parameter exponential distribution family. In the literature, several criteria have been prop…