29 citations · 54 across the 7 of their papers we have counts for
8 papers
A Unified Framework for Alternating Offline Model Training and Policy Learning
Shentao Yang, Shujian Zhang, Yihao Feng +1
In offline model-based reinforcement learning (offline MBRL), we learn a dynamic model from historically collected data, and subsequently utilize the learned model and fixed datase…
Incremental Few-shot Text Classification with Multi-round New Classes: Formulation, Dataset and System
Congying Xia, Wenpeng Yin, Yihao Feng +1
Text classification is usually studied by labeling natural language texts with relevant categories from a predefined set. In the real world, new classes might keep challenging the…
Non-asymptotic Confidence Intervals of Off-policy Evaluation: Primal and Dual Bounds
Yihao Feng, Ziyang Tang, Na Zhang +1
Off-policy evaluation (OPE) is the task of estimating the expected reward of a given policy based on offline data previously collected under different policies. Therefore, OPE is a…
Off-Policy Interval Estimation with Lipschitz Value Iteration
Ziyang Tang, Yihao Feng, Na Zhang +2
Off-policy evaluation provides an essential tool for evaluating the effects of different policies or treatments using only observed data. When applied to high-stakes scenarios such…
Accountable Off-Policy Evaluation With Kernel Bellman Statistics
Yihao Feng, Tongzheng Ren, Ziyang Tang +1
We consider off-policy evaluation (OPE), which evaluates the performance of a new policy from observed data collected from previous experiments, without requiring the execution of…
Doubly Robust Bias Reduction in Infinite Horizon Off-Policy Estimation
Ziyang Tang, Yihao Feng, Lihong Li +2
Infinite horizon off-policy policy evaluation is a highly challenging task due to the excessively large variance of typical importance sampling (IS) estimators. Recently, Liu et al…