Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
Hao Wang, Licheng Pan, Zhichao Chen +7
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…
cs.LG2026
Deep Time-series Forecasting Needs Kernelized Moment Balancing
Licheng Pan, Hao Wang, Haocheng Yang +7
Deep time-series forecasting can be formulated as a distribution balancing problem aimed at aligning the distribution of the forecasts and ground truths. According to Imbens' crite…
cs.LG2025
DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment
Hao Wang, Licheng Pan, Yuan Lu +7
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approac…