1 citations · 1 across the 3 of their papers we have counts for
1 paper · 1 filter
Shenzhi Wang, Qisen Yang, Jiawei Gao +6
Offline-to-online reinforcement learning (RL) is a training paradigm that combines pre-training on a pre-collected dataset with fine-tuning in an online environment. However, the i…