3 papers
cs.LG2026
Optimal and Efficient Contextual Combinatorial Semi-bandits with General Function Approximation
Hao Qin, Chicheng Zhang
We study the contextual combinatorial semi-bandit (CCSB) problem with general reward function approximation. At each round, the learner observes a context, selects a combinatorial…
cs.LG2025
Physics-Informed Parametric Bandits for Beam Alignment in mmWave Communications
Hao Qin, Thang Duong, Ming F. Li +1
In millimeter wave (mmWave) communications, beam alignment and tracking are crucial to combat the significant path loss. As scanning the entire directional space is inefficient, de…
cs.LG2025
Achieving adaptivity and optimality for multi-armed bandits using Exponential-Kullback Leibler Maillard Sampling
Hao Qin, Kwang-Sung Jun, Chicheng Zhang
We study the problem of -armed bandits with reward distributions belonging to a one-parameter exponential distribution family. In the literature, several criteria have been prop…