1 paper
Qinxun Bai, Yuxuan Han, Wei Xu +1
The actor-critic (AC) framework has achieved strong empirical success in off-policy reinforcement learning but suffers from the "moving target" problem, where the evaluated policy…