collaborators

7 papers

cs.LG2026

Learn to Match: Two-Sided Matching with Temporally Extended Feedback

Haijing Zong, Yancheng Liang, Boyang Zhou +1

Two-sided matching markets often involve information that unfolds over time through interviews, repeated interaction, learning, and separation. Existing matching models typically r…

cs.RO2026

E2HiL: Entropy-Guided Sample Selection for Efficient Real-World Human-in-the-Loop Reinforcement Learning

Haoyuan Deng, Yudong Lin, Yuanjiang Xue +5

Human-in-the-loop guidance has emerged as an effective approach for accelerating online reinforcement learning (RL) in real-world manipulation. However, existing human-in-the-loop…

math.ST2025

Local Asymptotic Normality for Multi-Armed Bandits

Ramon van den Akker, Bas J. M. Werker, Bo Zhou

Van den Akker, Werker, and Zhou (2025) showed that the limit experiment, in the sense of H\a'{a}jek-Le Cam, for (contextual) bandits whose arms' expected payoffs differ by $O(T^{-1…

econ.EM2025

Batched Adaptive Network Formation

Yan Xu, Bo Zhou

Networks are central to many economic and organizational applications, including workplace team formation, social platform recommendations, and classroom friendship development. In…

cs.RO2025

VMTS: Vision-Assisted Teacher-Student Reinforcement Learning for Multi-Terrain Locomotion in Bipedal Robots

Fu Chen, Rui Wan, Peidong Liu +2

Bipedal robots, due to their anthropomorphic design, offer substantial potential across various applications, yet their control is hindered by the complexity of their structure. Cu…

econ.EM2025

Valid Post-Contextual Bandit Inference

Ramon van den Akker, Bas J. M. Werker, Bo Zhou

We establish an asymptotic framework for the statistical analysis of the stochastic contextual multi-armed bandit problem (CMAB), which is widely employed in adaptively randomized…