4 citations · 9 across the 8 of their papers we have counts for
1 paper · 2 filters
Jiexin Wang, Eiji Uchibe
We introduce the ``soft Deep MaxPain'' (softDMP) algorithm, which integrates the optimization of long-term policy entropy into reward-punishment reinforcement learning objectives.…