2 citations · 2 across the 1 of their papers we have counts for
1 paper
Zihan Zhou, Wei Fu, Bingliang Zhang +1
We present Reward-Switching Policy Optimization (RSPO), a paradigm to discover diverse strategies in complex RL environments by iteratively finding novel policies that are both loc…