From the 1 of 6 linked papers with an AI index.
1 paper · 1 filter
Qijun Li, Zheng Fu, Qi Song +4
In complex continuous-control reinforcement learning tasks, multimodal optimal actions often coincide with uncertain, multimodal return distributions, making reliable value estimat…