1 paper
Seyeon Kim, Joonhun Lee, Namhoon Cho +2
Conventional uncertainty-aware temporal difference (TD) learning often models TD errors as zero-mean Gaussian. This assumption can miss the heavy-tailed and heteroscedastic residua…