2 citations · 2 across the 13 of their papers we have counts for
Showing 2026 · cs.LGShow all
3 papers · 2 filters
cs.LG2026
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
Hao Wang, Licheng Pan, Zhichao Chen +7
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…
cs.LG2026
Analyzing and Improving Diffusion Models for Time-Series Data Imputation: A Proximal Recursion Perspective
Zhichao Chen, Hao Wang, Fangyikang Wang +5
Diffusion models (DMs) have shown promise for Time-Series Data Imputation (TSDI); however, their performance remains inconsistent in complex scenarios. We attribute this to two pri…
cs.LG2026
Rethinking the Flow-Based Gradual Domain Adaptation: A Semi-Dual Optimal Transport Perspective
Zhichao Chen, Zhan Zhuang, Yunfei Teng +6
Gradual domain adaptation (GDA) aims to mitigate domain shift by progressively adapting models from the source domain to the target domain via intermediate domains. However, real i…