activity
20172026
most citedInformation-Theoretic Considerations in Batch Reinforcement Learning

49 citations · 77 across the 19 of their papers we have counts for

collaborators
Showing cs.LGShow all

26 papers · 1 filter

cs.LG2026

Q-MMR: Off-Policy Evaluation via Recursive Reweighting and Moment Matching

Xiang Li, Nan Jiang

We present a novel theoretical framework, Q-MMR, for off-policy evaluation in finite-horizon MDPs. Q-MMR learns a set of scalar weights, one for each data point, such that the rewe…

cs.LG2023

The Optimal Approximation Factors in Misspecified Off-Policy Value Function Estimation

Philip Amortila, Nan Jiang, Csaba Szepesvári

Theoretical guarantees in reinforcement learning (RL) are known to suffer multiplicative blow-up factors with respect to the misspecification error of function approximation. Yet,…

cs.LG2023

Offline Learning in Markov Games with General Function Approximation

Yuheng Zhang, Yu Bai, Nan Jiang

We study offline multi-agent reinforcement learning (RL) in Markov games, where the goal is to learn an approximate equilibrium -- such as Nash equilibrium and (Coarse) Correlated…

cs.LG2023

Reinforcement Learning in Low-Rank MDPs with Density Features

Audrey Huang, Jinglin Chen, Nan Jiang

MDPs with low-rank transitions -- that is, the transition matrix can be factored into the product of two matrices, left and right -- is a highly representative structure that enabl…

cs.LG2023★ 4 cited

Adversarial Model for Offline Reinforcement Learning

Mohak Bhardwaj, Tengyang Xie, Byron Boots +2

We propose a novel model-based offline Reinforcement Learning (RL) framework, called Adversarial Model for Offline Reinforcement Learning (ARMOR), which can robustly learn policies…

cs.LG2022★ 1 cited

Learning Markov Random Fields for Combinatorial Structures via Sampling through Lovász Local Lemma

Nan Jiang, Yi Gu, Yexiang Xue

Learning to generate complex combinatorial structures satisfying constraints will have transformative impacts in many application domains. However, it is beyond the capabilities of…