66 citations · 104 across the 7 of their papers we have counts for
18 papers · 1 filter
Exploiting Structure in Offline Multi-Agent RL: The Benefits of Low Interaction Rank
Wenhao Zhan, Scott Fujimoto, Zheqing Zhu +3
We study the problem of learning an approximate equilibrium in the offline multi-agent reinforcement learning (MARL) setting. We introduce a structural assumption -- the interactio…
Tractable Optimality in Episodic Latent MABs
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis +1
We consider a multi-armed bandit problem with latent contexts, where an agent interacts with the environment for an episode of time steps. Depending on the length of the ep…
Reward-Mixing MDPs with a Few Latent Contexts are Learnable
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis +1
We consider episodic reinforcement learning in reward-mixing Markov decision processes (RMMDPs): at the beginning of every episode nature randomly picks a latent reward model among…
Provable Reinforcement Learning with a Short-Term Memory
Yonathan Efroni, Chi Jin, Akshay Krishnamurthy +1
Real-world sequential decision making problems commonly involve partial observability, which requires the agent to maintain a memory of history in order to infer the latent states,…
Coordinated Attacks against Contextual Bandits: Fundamental Limits and Defense Mechanisms
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis +1
Motivated by online recommendation systems, we propose the problem of finding the optimal policy in multitask contextual bandits when a small fraction of tasks (users) are…
RL for Latent MDPs: Regret Guarantees and a Lower Bound
Jeongyeol Kwon, Yonathan Efroni, Constantine Caramanis +1
In this work, we consider the regret minimization problem for reinforcement learning in latent Markov Decision Processes (LMDP). In an LMDP, an MDP is randomly drawn from a set of…