3 papers
cs.LG2026
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
Jiesong Lian, Ruizhe Zhong, Zixiang Zhou +6
Post-training alignment of video generation models with human preferences is a critical goal. Developing effective Reward Models (RMs) for this process faces significant methodolog…
cs.LG2026
Euphonium: Steering Video Flow Matching via Process Reward Gradient Guided Stochastic Dynamics
Ruizhe Zhong, Jiesong Lian, Xiaoyue Mi +4
While online Reinforcement Learning has emerged as a crucial technique for aligning flow matching models with human preferences, current approaches are hindered by inefficient expl…
cs.CV2026
Video Generation Models Are Good Latent Reward Models
Xiaoyue Mi, Wenqing Yu, Jiesong Lian +9
Reward feedback learning (ReFL) has proven effective for aligning image generation with human preferences. However, its extension to video generation faces significant challenges.…