1 paper · 1 filter
Zhengpeng Xie, Jiahang Cao, Changwei Wang +5
In this paper, we argue that mutual distillation between reinforcement learning policies serves as an implicit regularization, preventing them from overfitting to irrelevant featur…