7 papers
Learn to Match: Two-Sided Matching with Temporally Extended Feedback
Haijing Zong, Yancheng Liang, Boyang Zhou +1
Two-sided matching markets often involve information that unfolds over time through interviews, repeated interaction, learning, and separation. Existing matching models typically r…
E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning
Haoyuan Deng, Yudong Lin, Yuanjiang Xue +5
Human-in-the-loop guidance has emerged as an effective approach for accelerating online reinforcement learning (RL) in real-world manipulation. However, existing human-in-the-loop…
Local Asymptotic Normality for Multi-Armed Bandits
Ramon van den Akker, Bas J. M. Werker, Bo Zhou
Van den Akker, Werker, and Zhou (2025) showed that the limit experiment, in the sense of H\a'{a}jek-Le Cam, for (contextual) bandits whose arms' expected payoffs differ by $O(T^{-1…
Batched Adaptive Network Formation
Yan Xu, Bo Zhou
Networks are central to many economic and organizational applications, including workplace team formation, social platform recommendations, and classroom friendship development. In…
VMTS: Vision-Assisted Teacher-Student Reinforcement Learning for Multi-Terrain Locomotion in Bipedal Robots
Fu Chen, Rui Wan, Peidong Liu +2
Bipedal robots, due to their anthropomorphic design, offer substantial potential across various applications, yet their control is hindered by the complexity of their structure. Cu…
Valid Post-Contextual Bandit Inference
Ramon van den Akker, Bas J. M. Werker, Bo Zhou
We establish an asymptotic framework for the statistical analysis of the stochastic contextual multi-armed bandit problem (CMAB), which is widely employed in adaptively randomized…