2 citations · 2 across the 4 of their papers we have counts for
1 paper · 1 filter
Rajdeep Dutta, Qincheng Wang, Ankur Singh +3
This paper presents a novel RL algorithm, S-REINFORCE, which is designed to generate interpretable policies for dynamic decision-making tasks. The proposed algorithm leverages two…