collaborators

6 papers

cs.CV2026

SARA: Semantically Adaptive Relational Alignment for Video Diffusion Models

Jiesong Lian, Zixiang Zhou, Ruizhe Zhong +6

Recent video diffusion models (VDMs) synthesize visually convincing clips, yet still drop entities, mis-bind attributes, and weaken the interactions specified in the prompt. Repres…

cs.LG2026

Asymmetric On-Policy Distillation: Bridging Exploitation and Imitation at the Token Level

Nan Jia, Haojin Yang, Xing Ma +6

On-policy distillation (OPD) trains a student on its own trajectories with token-level teacher feedback and often outperforms off-policy distillation and standard reinforcement lea…

cs.LG2026

SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models

Jiesong Lian, Ruizhe Zhong, Zixiang Zhou +6

Post-training alignment of video generation models with human preferences is a critical goal. Developing effective Reward Models (RMs) for this process faces significant methodolog…

cs.LG2026

Euphonium: Steering Video Flow Matching via Process Reward Gradient Guided Stochastic Dynamics

Ruizhe Zhong, Jiesong Lian, Xiaoyue Mi +4

While online Reinforcement Learning has emerged as a crucial technique for aligning flow matching models with human preferences, current approaches are hindered by inefficient expl…

cs.CV2026

Video Generation Models Are Good Latent Reward Models

Xiaoyue Mi, Wenqing Yu, Jiesong Lian +9

Reward feedback learning (ReFL) has proven effective for aligning image generation with human preferences. However, its extension to video generation faces significant challenges.…

cs.GT2026

Fusion-PSRO: Nash Policy Fusion for Policy Space Response Oracles

Jiesong Lian, Yucong Huang, Chengdong Ma +4

For solving zero-sum games involving non-transitivity, a useful approach is to maintain a policy population to approximate the Nash Equilibrium (NE). Previous studies have shown th…