6 citations · 6 across the 2 of their papers we have counts for
1 paper · 1 filter
Liting Chen, Jie Yan, Zhengdao Shao +5
Offline reinforcement learning faces a significant challenge of value over-estimation due to the distributional drift between the dataset and the current learned policy, leading to…