4 papers
Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes
Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi +2
Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The or…
Information-Directed Sampling for Causal Bandits
Muhammad Qasim Elahi, Murat Kocaoglu, Mahsa Ghasemi
Causal bandits exploit structural relationships among variables to share information across interventions and accelerate the identification of high-reward decisions. In many applic…
Joint MDPs and Reinforcement Learning in Coupled-Dynamics Environments
Ege C. Kaya, Mahsa Ghasemi, Abolfazl Hashemi
Many distributional quantities in reinforcement learning are intrinsically joint across actions, including distributions of gaps and probabilities of superiority. However, the clas…
Partial Structure Discovery is Sufficient for No-regret Learning in Causal Bandits
Muhammad Qasim Elahi, Mahsa Ghasemi, Murat Kocaoglu
Causal knowledge about the relationships among decision variables and a reward variable in a bandit setting can accelerate the learning of an optimal decision. Current works often…