3 citations · 11 across the 5 of their papers we have counts for
4 papers · 1 filter
Off-Policy Evaluation for Human Feedback
Qitong Gao, Ge Gao, Juncheng Dong +3
Off-policy evaluation (OPE) is important for closing the gap between offline training and evaluation of reinforcement learning (RL), by estimating performance and/or rank of target…
Robust Reinforcement Learning through Efficient Adversarial Herding
Juncheng Dong, Hao-Lun Hsu, Qitong Gao +2
Although reinforcement learning (RL) is considered the gold standard for policy design, it may not always provide a robust solution in various scenarios. This can result in severe…
PASTA: Pessimistic Assortment Optimization
Juncheng Dong, Weibin Mo, Zhengling Qi +3
We consider a class of assortment optimization problems in an offline data-driven setting. A firm does not know the underlying customer choice model but has access to an offline da…
Domain Adaptation via Rebalanced Sub-domain Alignment
Yiling Liu, Juncheng Dong, Ziyang Jiang +5
Unsupervised domain adaptation (UDA) is a technique used to transfer knowledge from a labeled source domain to a different but related unlabeled target domain. While many UDA metho…