activity
20172021
most citedInformation-Theoretic Considerations in Batch Reinforcement Learning

49 citations · 70 across the 6 of their papers we have counts for

collaborators

19 papers

cs.LG20213 cited

Towards Hyperparameter-free Policy Selection for Offline Reinforcement Learning

Siyuan Zhang, Nan Jiang

How to select between policies and value functions produced by different training algorithms in offline reinforcement learning (RL) -- which is crucial for hyperpa-rameter tuning -…

cs.LG2021

Minimax Model Learning

Cameron Voloshin, Nan Jiang, Yisong Yue

We present a novel off-policy loss function for learning a transition model in model-based reinforcement learning. Notably, our loss is derived from the off-policy policy evaluatio…

cs.LG2021

On Query-efficient Planning in MDPs under Linear Realizability of the Optimal State-value Function

Gellért Weisz, Philip Amortila, Barnabás Janzer +3

We consider local planning in fixed-horizon MDPs with a generative model under the assumption that the optimal value function lies close to the span of a feature map. The generativ…

math.OC2020

ALSO-X and ALSO-X+: Better Convex Approximations for Chance Constrained Programs

Nan Jiang, Weijun Xie

In a chance constrained program (CCP), the decision-makers aim to seek the best decision whose probability of violating the uncertainty constraints is within the prespecified risk…

cs.LG20209 cited

A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting

Philip Amortila, Nan Jiang, Tengyang Xie

Recently, Wang et al. (2020) showed a highly intriguing hardness result for batch reinforcement learning (RL) with linearly realizable value function and good feature coverage in t…

cs.LG2020

Improved Worst-Case Regret Bounds for Randomized Least-Squares Value Iteration

Priyank Agrawal, Jinglin Chen, Nan Jiang

This paper studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping varia…