10 papers
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
Hao Wang, Haocheng Yang, Licheng Pan +7
Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…
Deep Autocorrelation Modeling for Time-Series Forecasting: Progress and Prospects
Hao Wang, Licheng Pan, Qingsong Wen +12
Autocorrelation is a defining characteristic of time-series data, where each observation is statistically dependent on its predecessors. In the context of deep time-series forecast…
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
Hao Wang, Licheng Pan, Zhichao Chen +7
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…
Analyzing and Improving Diffusion Models for Time-Series Data Imputation: A Proximal Recursion Perspective
Zhichao Chen, Hao Wang, Fangyikang Wang +5
Diffusion models (DMs) have shown promise for Time-Series Data Imputation (TSDI); however, their performance remains inconsistent in complex scenarios. We attribute this to two pri…
Deep Time-series Forecasting Needs Kernelized Moment Balancing
Licheng Pan, Hao Wang, Haocheng Yang +7
Deep time-series forecasting can be formulated as a distribution balancing problem aimed at aligning the distribution of the forecasts and ground truths. According to Imbens' crite…
Quadratic Direct Forecast for Training Multi-Step Time-Series Forecast Models
Hao Wang, Licheng Pan, Yuan Lu +7
The design of training objective is central to training time-series forecasting models. Existing training objectives such as mean squared error mostly treat each future step as an…