3 citations · 3 across the 1 of their papers we have counts for
4 papers
WeightNet: Revisiting the Design Space of Weight Networks
Ningning Ma, Xiangyu Zhang, Jiawei Huang +1
We present a conceptually simple, flexible and effective framework for weight generating networks. Our approach is general that unifies two current distinct and extremely effective…
Minimax Value Interval for Off-Policy Evaluation and Policy Optimization
Nan Jiang, Jiawei Huang
We study minimax methods for off-policy evaluation (OPE) using value functions and marginalized importance weights. Despite that they hold promises of overcoming the exponential va…
Minimax Weight and Q-Function Learning for Off-Policy Evaluation
Masatoshi Uehara, Jiawei Huang, Nan Jiang
We provide theoretical investigations into off-policy evaluation in reinforcement learning using function approximators for (marginalized) importance weights and value functions. O…
From Importance Sampling to Doubly Robust Policy Gradient
Jiawei Huang, Nan Jiang
We show that on-policy policy gradient (PG) and its variance reduction variants can be derived by taking finite difference of function evaluations supplied by estimators from the i…