49 citations · 70 across the 6 of their papers we have counts for
19 papers
Towards Hyperparameter-free Policy Selection for Offline Reinforcement Learning
Siyuan Zhang, Nan Jiang
How to select between policies and value functions produced by different training algorithms in offline reinforcement learning (RL) -- which is crucial for hyperpa-rameter tuning -…
Minimax Model Learning
Cameron Voloshin, Nan Jiang, Yisong Yue
We present a novel off-policy loss function for learning a transition model in model-based reinforcement learning. Notably, our loss is derived from the off-policy policy evaluatio…
On Query-efficient Planning in MDPs under Linear Realizability of the Optimal State-value Function
Gellért Weisz, Philip Amortila, Barnabás Janzer +3
We consider local planning in fixed-horizon MDPs with a generative model under the assumption that the optimal value function lies close to the span of a feature map. The generativ…
ALSO-X and ALSO-X+: Better Convex Approximations for Chance Constrained Programs
Nan Jiang, Weijun Xie
In a chance constrained program (CCP), the decision-makers aim to seek the best decision whose probability of violating the uncertainty constraints is within the prespecified risk…
A Variant of the Wang-Foster-Kakade Lower Bound for the Discounted Setting
Philip Amortila, Nan Jiang, Tengyang Xie
Recently, Wang et al. (2020) showed a highly intriguing hardness result for batch reinforcement learning (RL) with linearly realizable value function and good feature coverage in t…
Improved Worst-Case Regret Bounds for Randomized Least-Squares Value Iteration
Priyank Agrawal, Jinglin Chen, Nan Jiang
This paper studies regret minimization with randomized value functions in reinforcement learning. In tabular finite-horizon Markov Decision Processes, we introduce a clipping varia…