8 papers
Rubrics on Trial: Evolving Rubrics from a Single Query via Synthetic Pairwise Evidence
Haocheng Yang, Licheng Pan, Xiaoxi Li +5
The paper proposes a query‑only method that automatically creates and validates fine‑grained rubrics for evaluating large language models by using synthetic rubric‑conditioned resp…
DistDF: Time-Series Forecasting Needs Joint-Distribution Wasserstein Alignment
Hao Wang, Licheng Pan, Yuan Lu +7
Training time-series forecasting models requires aligning the conditional distribution of model forecasts with that of the label sequence. The standard direct forecast (DF) approac…
MuSEAgent: A Multimodal Reasoning Agent with Stateful Experiences
Shijian Wang, Jiarui Jin, Runhao Fu +11
Research agents have recently achieved significant progress in information seeking and synthesis across heterogeneous textual and visual sources. In this paper, we introduce MuSEAg…
ImplicitRM: Unbiased Reward Modeling from Implicit Preference Data for LLM alignment
Hao Wang, Haocheng Yang, Licheng Pan +7
Reward modeling represents a long-standing challenge in reinforcement learning from human feedback (RLHF) for aligning language models. Current reward modeling is heavily contingen…
Deep Autocorrelation Modeling for Time-Series Forecasting: Progress and Prospects
Hao Wang, Licheng Pan, Qingsong Wen +12
Autocorrelation is a defining characteristic of time-series data, where each observation is statistically dependent on its predecessors. In the context of deep time-series forecast…
CausalRM: Causal-Theoretic Reward Modeling for RLHF from Observational User Feedbacks
Hao Wang, Licheng Pan, Zhichao Chen +7
Despite the success of reinforcement learning from human feedback (RLHF) in aligning language models, current reward modeling heavily relies on experimental feedback data collected…