4 papers
Variance-aware Reward Modeling with Anchor Guidance
Shuxing Fang, Ruijian Han, Liangyu Zhang +1
Standard Bradley--Terry (BT) reward models are limited when human preferences are pluralistic. Although soft preference labels preserve disagreement information, BT can only expres…
Deep Ranking with Heterogeneous Effects
Yuanhang Luo, Shuxing Fang, Ruijian Han +1
Classical latent-score ranking models often fail to distinguish objects' intrinsic scores from contextual effects, which are typically nonlinear and can dominate the observed outco…
Recent advances in the Bradley--Terry Model: theory, algorithms, and applications
Shuxing Fang, Ruijian Han, Yuanhang Luo +1
This article surveys recent progress in the Bradley-Terry (BT) model and its extensions. We focus on the statistical and computational aspects, with emphasis on the regime in which…
Two-way Deconfounder for Off-policy Evaluation in Causal Reinforcement Learning
Shuguang Yu, Shuxing Fang, Ruixin Peng +3
This paper studies off-policy evaluation (OPE) in the presence of unmeasured confounders. Inspired by the two-way fixed effects regression model widely used in the panel data liter…