49 citations · 77 across the 19 of their papers we have counts for
26 papers · 1 filter
Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching
Xiang Li, Nan Jiang
We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the rewe…
The Optimal Approximation Factors in Misspecified Off-Policy Value Function Estimation
Philip Amortila, Nan Jiang, Csaba Szepesvári
Theoretical guarantees in reinforcement learning (RL) are known to suffer multiplicative blow-up factors with respect to the misspecification error of function approximation. Yet,…
Offline Learning in Markov Games with General Function Approximation
Yuheng Zhang, Yu Bai, Nan Jiang
We study offline multi-agent reinforcement learning (RL) in Markov games, where the goal is to learn an approximate equilibrium -- such as Nash equilibrium and (Coarse) Correlated…
Reinforcement Learning in Low-Rank MDPs with Density Features
Audrey Huang, Jinglin Chen, Nan Jiang
MDPs with low-rank transitions -- that is, the transition matrix can be factored into the product of two matrices, left and right -- is a highly representative structure that enabl…
Adversarial Model for Offline Reinforcement Learning
Mohak Bhardwaj, Tengyang Xie, Byron Boots +2
We propose a novel model-based offline Reinforcement Learning (RL) framework, called Adversarial Model for Offline Reinforcement Learning (ARMOR), which can robustly learn policies…
Learning Markov Random Fields for Combinatorial Structures via Sampling through Lovász Local Lemma
Nan Jiang, Yi Gu, Yexiang Xue
Learning to generate complex combinatorial structures satisfying constraints will have transformative impacts in many application domains. However, it is beyond the capabilities of…