1k citations · 1.5k across the 7 of their papers we have counts for
1 paper · 1 filter
Steven Hansen, Will Dabney, Andre Barreto +3
It has been established that diverse behaviors spanning the controllable subspace of an Markov decision process can be trained by rewarding a policy for being distinguishable from…