96 citations · 96 across the 1 of their papers we have counts for
1 paper
Elad Hazan, Sham M. Kakade, Karan Singh +1
Suppose an agent is in a (possibly unknown) Markov Decision Process in the absence of a reward signal, what might we hope that an agent can efficiently learn to do? This work studi…