3 citations · 3 across the 1 of their papers we have counts for
1 paper · 1 filter
Yang Cao, Zheng Wen, Branislav Kveton +1
Multi-armed bandit (MAB) is a class of online learning problems where a learning agent aims to maximize its expected cumulative reward while repeatedly selecting to pull arms with…