7 citations · 31 across the 13 of their papers we have counts for
4 papers · 1 filter
Off-Policy Evaluation for Human Feedback
Qitong Gao, Ge Gao, Juncheng Dong +3
Off-policy evaluation (OPE) is important for closing the gap between offline training and evaluation of reinforcement learning (RL), by estimating performance and/or rank of target…
An Offline Time-aware Apprenticeship Learning Framework for Evolving Reward Functions
Xi Yang, Ge Gao, Min Chi
Apprenticeship learning (AL) is a process of inducing effective decision-making policies via observing and imitating experts' demonstrations. Most existing AL approaches, however,…
HOPE: Human-Centric Off-Policy Evaluation for E-Learning and Healthcare
Ge Gao, Song Ju, Markel Sanz Ausin +1
Reinforcement learning (RL) has been extensively researched for enhancing human-environment interactions in various human-centric tasks, including e-learning and healthcare. Since…
Variational Latent Branching Model for Off-Policy Evaluation
Qitong Gao, Ge Gao, Min Chi +1
Model-based methods have recently shown great potential for off-policy evaluation (OPE); offline trajectories induced by behavioral policies are fitted to transitions of Markov dec…