From the 2 of 24 linked papers with an AI index.
24 papers
Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
Rui Li, Yuanzhi Liang, Ke Hao +4
Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. Howeve…
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Rui Li, Yuanzhi Liang, Ziqi Ni +3
The paper proposes TaRoS, a framework that redesigns reward signals for video generation using GRPO to avoid reward hacking and saturation, improving visual fidelity, motion cohere…
From World Action Models to Embodied Brains: A Roadmap for Open-World Physical Intelligence
Yuanzhi Liang, Xufeng Zhan, Haibin Huang +2
The paper outlines a roadmap for building open‑world physical intelligence by integrating World Action Models with an "embodied brain" architecture that unifies multimodal context,…
VideoWeave: Unlocking Geometric Consistency in Video Generation via Joint Geometry-Video Modeling
Xunzhi Xiang, Zixuan Duan, Yabo Chen +8
Large-scale video diffusion models often fail to preserve 3D structure over time, causing geometric drift and implausible motion under viewpoint changes. Existing methods usually e…
Reinforcing Few-step Generators via Reward-Tilted Distribution Matching
Yushi Huang, Xiangxin Zhou, Ruoyu Wang +3
Recent advances in few-step diffusion distillation have enabled efficient image generation, yet aligning these models with human preferences remains challenging. We propose Reward-…
Full-4D: Generating Full-Scope 4D Scenes from a Single-View Video
Tingxi Chen, Ke Hao, Yabo Chen +6
Generating 4D scenes from a single-view video is inherently ill-posed: a single viewpoint lacks the information needed to recover a complete, dynamic scene with full coverage. Exis…