2 papers
cs.LG2026
ROAD: Adaptive Data Mixing for Offline-to-Online Reinforcement Learning via Bi-Level Optimization
Letian Yang, Xu Liu, Yiqiang Lu +3
Offline-to-online reinforcement learning harnesses the stability of offline pretraining and the flexibility of online fine-tuning. A key challenge lies in the non-stationary distri…
cs.LG2025
Stochastically Constrained Best Arm Identification with Thompson Sampling
Le Yang, Siyang Gao, Cheng Li +1
We consider the problem of the best arm identification in the presence of stochastic constraints, where there is a finite number of arms associated with multiple performance measur…