From the 1 of 23 linked papers with an AI index.
2 citations · 3 across the 11 of their papers we have counts for
12 papers · 1 filter
Optimal Transport for LLM Reward Modeling from Noisy Preference
Licheng Pan, Haochen Yang, Haoxuan Li +8
Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…
DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment
Hao Wang, Licheng Pan, Yuan Lu +7
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approac…
Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation
Hao Wang, Zhichao Chen, Zhaoran Liu +3
Heterogeneous treatment effect (HTE) estimation from observational data poses significant challenges due to treatment selection bias. Existing methods address this bias by minimizi…
Entire Space Counterfactual Learning for Reliable Content Recommendations
Hao Wang, Zhichao Chen, Zhaoran Liu +4
Post-click conversion rate (CVR) estimation is a fundamental task in developing effective recommender systems, yet it faces challenges from data sparsity and sample selection bias.…
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
Hao Wang, Licheng Pan, Zhichao Chen +7
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…
Deep Time-series Forecasting Needs Kernelized Moment Balancing
Licheng Pan, Hao Wang, Haocheng Yang +7
Deep time-series forecasting can be formulated as a distribution balancing problem aimed at aligning the distribution of the forecasts and ground truths. According to Imbens' crite…