42 citations · 42 across the 1 of their papers we have counts for
1 paper
Mohammad Gheshlaghi Azar, Remi Munos, Bert Kappen
We consider the problem of learning the optimal action-value function in the discounted-reward Markov decision processes (MDPs). We prove a new PAC bound on the sample-complexity o…