2 papers
cs.LG2024
Online Policy Learning from Offline Preferences
Guoxi Zhang, Han Bao, Hisashi Kashima
In preference-based reinforcement learning (PbRL), a reward function is learned from a type of human feedback called preference. To expedite preference collection, recent works hav…
cs.LG2023
Estimating Treatment Effects Under Heterogeneous Interference
Xiaofeng Lin, Guoxi Zhang, Xiaotian Lu +3
Treatment effect estimation can assist in effective decision-making in e-commerce, medicine, and education. One popular application of this estimation lies in the prediction of the…