1 paper
Bryan Chan, Anson Leung, James Bergstra
Offline-to-online reinforcement learning (O2O RL) aims to obtain a continually improving policy as it interacts with the environment, while ensuring the initial policy behaviour is…