5 papers
Local Asymptotic Normality for Multi-Armed Bandits
Ramon van den Akker, Bas J. M. Werker, Bo Zhou
Van den Akker, Werker, and Zhou (2025) showed that the limit experiment, in the sense of H\a'{a}jek-Le Cam, for (contextual) bandits whose arms' expected payoffs differ by $O(T^{-1…
Batched Adaptive Network Formation
Yan Xu, Bo Zhou
Networks are central to many economic and organizational applications, including workplace team formation, social platform recommendations, and classroom friendship development. In…
Valid Post-Contextual Bandit Inference
Ramon van den Akker, Bas J. M. Werker, Bo Zhou
We establish an asymptotic framework for the statistical analysis of the stochastic contextual multi-armed bandit problem (CMAB), which is widely employed in adaptively randomized…
VMTS: Vision-Assisted Teacher-Student Reinforcement Learning for Multi-Terrain Locomotion in Bipedal Robots
Fu Chen, Rui Wan, Peidong Liu +2
Bipedal robots, due to their anthropomorphic design, offer substantial potential across various applications, yet their control is hindered by the complexity of their structure. Cu…
Survey on Strategic Mining in Blockchain: A Reinforcement Learning Approach
Jichen Li, Lijia Xie, Hanting Huang +5
Strategic mining attacks, such as selfish mining, exploit blockchain consensus protocols by deviating from honest behavior to maximize rewards. Markov Decision Process (MDP) analys…