works on

From the 1 of 23 linked papers with an AI index.

activity
20242026
most citedEntire Space Counterfactual Learning for Reliable Content Recommendations

2 citations · 3 across the 11 of their papers we have counts for

collaborators
Showing cs.LGShow all

12 papers · 1 filter

cs.LG2026

Optimal Transport for LLM Reward Modeling from Noisy Preference

Licheng Pan, Haochen Yang, Haoxuan Li +8

Reward models are fundamental to Reinforcement Learning from Human Feedback (RLHF), yet real-world datasets are inevitably corrupted by noisy preference. Conventional training obje…

cs.LG2026

DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment

Hao Wang, Licheng Pan, Yuan Lu +7

Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approac…

cs.LG2026

Proximity Matters: Local Proximity Enhanced Balancing for Treatment Effect Estimation

Hao Wang, Zhichao Chen, Zhaoran Liu +3

Heterogeneous treatment effect (HTE) estimation from observational data poses significant challenges due to treatment selection bias. Existing methods address this bias by minimizi…

cs.LG20262 cited

Entire Space Counterfactual Learning for Reliable Content Recommendations

Hao Wang, Zhichao Chen, Zhaoran Liu +4

Post-click conversion rate (CVR) estimation is a fundamental task in developing effective recommender systems, yet it faces challenges from data sparsity and sample selection bias.…

cs.LG2026

CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks

Hao Wang, Licheng Pan, Zhichao Chen +7

Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…

cs.LG2026

Deep Time-series Forecasting Needs Kernelized Moment Balancing

Licheng Pan, Hao Wang, Haocheng Yang +7

Deep time-series forecasting can be formulated as a distribution balancing problem aimed at aligning the distribution of the forecasts and ground truths. According to Imbens' crite…