2 citations · 2 across the 1 of their papers we have counts for
1 paper
Adrian Rivera Cardoso, He Wang, Huan Xu
We consider Markov Decision Processes (MDPs) where the rewards are unknown and may change in an adversarial manner. We provide an algorithm that achieves state-of-the-art regret bo…