8 citations · 22 across the 12 of their papers we have counts for
3 papers · 1 filter
Tractable Optimality in Episodic Latent MABs
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis +1
We consider a multi-armed bandit problem with latent contexts, where an agent interacts with the environment for an episode of time steps. Depending on the length of the ep…
Reward-Mixing MDPs with a Few Latent Contexts are Learnable
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis +1
We consider episodic reinforcement learning in reward-mixing Markov decision processes (RMMDPs): at the beginning of every episode nature randomly picks a latent reward model among…
Coordinated Attacks against Contextual Bandits: Fundamental Limits and Defense Mechanisms
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis +1
Motivated by online recommendation systems, we propose the problem of finding the optimal policy in multitask contextual bandits when a small fraction of tasks (users) are…