1 paper
Denis Tarasov, Kirill Brilliantov, Dmitrii Kharlapenko
In deep Reinforcement Learning (RL), value functions are typically approximated using deep neural networks and trained via mean squared error regression objectives to fit the true…