1 paper · 1 filter
Zhi Wang, Chicheng Zhang, Manish Kumar Singh +2
In many real-world applications, multiple agents seek to learn how to perform highly related yet slightly different tasks in an online bandit learning protocol. We formulate this p…