1 paper
Jing Fu, Bill Moran, José Niño-Mora
We study a system with finitely many groups of multi-action bandit processes, each of which is a Markov decision process (MDP) with finite state and action spaces and potentially d…