activity
20172022
most citedLearning to Draw Samples with Amortized Stein Variational Gradient Descent

29 citations · 54 across the 7 of their papers we have counts for

collaborators

8 papers

cs.LG20222 cited

A Unified Framework for Alternating Offline Model Training and Policy Learning

Shentao Yang, Shujian Zhang, Yihao Feng +1

In offline model-based reinforcement learning (offline MBRL), we learn a dynamic model from historically collected data, and subsequently utilize the learned model and fixed datase…

cs.CL20213 cited

Incremental Few-shot Text Classification with Multi-round New Classes: Formulation, Dataset and System

Congying Xia, Wenpeng Yin, Yihao Feng +1

Text classification is usually studied by labeling natural language texts with relevant categories from a predefined set. In the real world, new classes might keep challenging the…

cs.LG20211 cited

Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds

Yihao Feng, Ziyang Tang, Na Zhang +1

Off-policy evaluation (OPE) is the task of estimating the expected reward of a given policy based on offline data previously collected under different policies. Therefore, OPE is a…

cs.LG2020

Off-Policy Interval Estimation with Lipschitz Value Iteration

Ziyang Tang, Yihao Feng, Na Zhang +2

Off-policy evaluation provides an essential tool for evaluating the effects of different policies or treatments using only observed data. When applied to high-stakes scenarios such…

cs.LG20207 cited

Accountable Off-Policy Evaluation With Kernel Bellman Statistics

Yihao Feng, Tongzheng Ren, Ziyang Tang +1

We consider off-policy evaluation (OPE), which evaluates the performance of a new policy from observed data collected from previous experiments, without requiring the execution of…

cs.LG201912 cited

Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation

Ziyang Tang, Yihao Feng, Lihong Li +2

Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al…