2 citations · 2 across the 1 of their papers we have counts for
1 paper · 1 filter
Mohammadi Zaki, Avinash Mohan, Aditya Gopalan +1
We consider an improper reinforcement learning setting where a learner is given M base controllers for an unknown Markov decision process, and wishes to combine them optimally to…