4 papers
Lookahead Tree-Based Rollouts for Enhanced Trajectory-Level Exploration in Reinforcement Learning with Verifiable Rewards
Shangyu Xing, Siyuan Wang, Chenyuan Yang +2
Reinforcement Learning with Verifiable Rewards (RLVR), particularly with algorithms like Group Relative Policy Optimization (GRPO), has proven highly effective in enhancing the rea…
Dynamic Decision-Making under Model Misspecification: A Stochastic Stability Approach
Xinyu Dai, Daniel Chen, Yian Qian
Dynamic decision-making under model uncertainty is central to many economic environments, yet existing bandit and reinforcement learning algorithms rely on the assumption of correc…
Sloan Digital Sky Survey-V: Pioneering Panoptic Spectroscopy
Juna A. Kollmeier, Hans-Walter Rix, Conny Aerts +217
The Sloan Digital Sky Survey-V (SDSS-V) is pioneering panoptic spectroscopy: it is the first all-sky, multi-epoch, optical-to-infrared spectroscopic survey. SDSS-V is mapping the s…
Dynamic Decision-Making under Model Misspecification
Xinyu Dai
In this study, I investigate the dynamic decision problem with a finite parameter space when the functional form of conditional expected rewards is misspecified. Traditional algori…