From the 2 of 21 linked papers with an AI index.
21 papers
Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
Rui Li, Yuanzhi Liang, Ke Hao +4
Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. Howeve…
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Rui Li, Yuanzhi Liang, Ziqi Ni +3
The paper proposes TaRoS, a framework that redesigns reward signals for video generation using GRPO to avoid reward hacking and saturation, improving visual fidelity, motion cohere…
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Yuanzhi Liang, Xufeng Zhan, Haibin Huang +2
The paper outlines a roadmap for building open‑world physical intelligence by integrating World Action Models with an "embodied brain" architecture that unifies multimodal context,…
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
Ziqi Ni, Yuanzhi Liang, Rui Li +4
Reinforcement learning (RL) has become a powerful tool for post-training visual generative models, with Group Relative Policy Optimization (GRPO) increasingly used to align generat…
Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation
Rui Li, Ke Hao, Yuanzhi Liang +4
Reinforcement learning, particularly Group Relative Policy Optimization (GRPO), has emerged as an effective framework for post-training visual generative models with human preferen…
Reward-Aware Trajectory Shaping for Few-step Visual Generation
Rui Li, Bingyu Li, Yuanzhi Liang +3
Achieving high-fidelity generation in extremely few sampling steps has long been a central goal of generative modeling. Existing approaches largely rely on distillation-based frame…