1 paper
Jikai Jin, Kenneth Hung, Sanath Kumar Krishnamurthy +2
We study bandits whose rewards depend on an unobserved Markov state that evolves independently of the learner's actions. The optimal arm can change even though the learner observes…