Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Payel Bhattacharjee, Osvaldo Simeone, Ravi Tandon
Reward modeling is central to RLHF, RLAIF, and PPO-based alignment, but its reliability is often limited by scarce and heterogeneous human preference data. In this paper, we introd…
cs.LG2026
Prediction-Powered Risk Monitoring of Deployed Models for Detecting Harmful Distribution Shifts
Guangyi Zhang, Yunlong Cai, Guanding Yu +1
We study the problem of monitoring model performance in dynamic environments where labeled data are limited. To this end, we propose prediction-powered risk monitoring (PPRM), a se…