2 papers
cs.LG2022
Reward-Biased Maximum Likelihood Estimation for Neural Contextual Bandits
Yu-Heng Hung, Ping-Chun Hsieh
Reward-biased maximum likelihood estimation (RBMLE) is a classic principle in the adaptive control literature for tackling explore-exploit trade-offs. This paper studies the stocha…
cs.LG2020
Reward-Biased Maximum Likelihood Estimation for Linear Stochastic Bandits
Yu-Heng Hung, Ping-Chun Hsieh, Xi Liu +1
Modifying the reward-biased maximum likelihood method originally proposed in the adaptive control literature, we propose novel learning algorithms to handle the explore-exploit tra…