1 citations · 1 across the 1 of their papers we have counts for
1 paper
Shenzhi Wang, Qisen Yang, Jiawei Gao +6
Offline-to-online reinforcement learning (RL) is a training paradigm that combines pre-training on a pre-collected dataset with fine-tuning in an online environment. However, the i…