5 papers · 1 filter
Provably Efficient Reinforcement Learning in Continuous-Time Episodic MDPs with Poisson Decision Epochs
Kenny Guo, Valentio Iverson, Sahan Wijetunga +1
Many real-world reinforcement learning (RL) problems evolve in continuous time, where decisions occur at irregular, event-driven intervals rather than at fixed discrete steps. We s…
The Price of Decentralization in Top- Arm Identification
Larissa Xu, Jasmine Nguyen, William Chang
Cooperative teams often need to agree on the best few options rather than simply accumulate reward, and they must do so while each member sees only a fragment of the team's collect…
Decentralized Multi-Player Q-Learning in Episodic Markov Decision Processes with Information Asymmetry
Larissa Xu, King Bi, William Chang
We study decentralized multi-player reinforcement learning in episodic tabular Markov decision processes (MDPs) under three forms of information asymmetry: (A) unobserved actions w…
DCM Bandits: Multiplayer Information Asymmetric Cascading Bandits for Multiple Clicks
Andy Wang, Charlton Shih, William Chang
In this work, we extend the Dependent Click Model (DCM) Bandits to a multiplayer information-asymmetric setting, where multiple agents interact with a shared ranked list and may ob…
Coordinating the Unknown Lipschitz Constant in Multiplayer Bandits
Ricardo Parada, Chenzhang Zhao, William Chang
Motivated by decentralized applications, we study cooperative multi-agent bandits in continuous (Lipschitz) action spaces when the Lipschitz constant is unknown. We consider three…