5 citations · 5 across the 3 of their papers we have counts for
1 paper · 1 filter
Juan Cervino, Harshat Kumar, Alejandro Ribeiro
We consider the problem of finding a policy that maximizes an expected reward throughout the trajectory of an agent that interacts with an unknown environment. Frequently denoted R…