From the 1 of 4 linked papers with an AI index.
4 papers
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Rui Li, Yuanzhi Liang, Ziqi Ni +3
The paper proposes TaRoS, a framework that redesigns reward signals for video generation using GRPO to avoid reward hacking and saturation, improving visual fidelity, motion cohere…
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
Ziqi Ni, Yuanzhi Liang, Rui Li +4
Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generat…
TeleBoost: A Systematic Alignment Framework for High-Fidelity, Controllable, and Robust Video Generation
Yuanzhi Liang, Xuan'er Wu, Yirui Liu +12
Post-training is the decisive step for converting a pretrained video generator into a production-oriented model that is instruction-following, controllable, and robust over long te…
Integrating Reinforcement Learning with Visual Generative Models: Foundations and Advances
Yuanzhi Liang, Yijie Fang, Ke Hao +4
Generative models have made significant progress in synthesizing visual content, including images, videos, and 3D/4D structures. However, they are typically trained with surrogate…