32 papers
Sample-Adaptive Latent Rewards for Uncertainty-Guided Diffusion Post-Training
Rui Li, Yuanzhi Liang, Ke Hao +4
Latent reward models can supervise visual diffusion models without decoding intermediate states into pixel space. This makes alignment with human preferences more efficient. Howeve…
CineWeaver: Training-Free Reference-Controllable Multi-Shot Long Video Generation for Cinematic Storytelling
Yuyang Huang, Yabo Chen, Wenrui Dai +6
CineWeaver introduces a training-free method that modifies pretrained video diffusion models to generate long, multi-shot cinematic videos with fine-grained reference control and c…
MMOE: Modernizing Diffusion Transformers with Efficient Expert Design
Yanhao Jia, Jiepeng Wang, Haibin Huang +3
Modern large language models scale successfully by pairing capacity growth with efficiency, keeping per-token and deployment costs under control as capacity grows. AIGC Foundation…
ShotPlan: Cinematic Video Generation with Learnable Planning Token
Su Guo, Guangce Liu, Haosen Yang +7
Current video generation models achieve impressive results in single-shot generation, yet remain limited in cinematic video generation, where coherent narratives and effective mult…
Generative Transmission: Rethinking Computation, Bandwidth, and Memory in Communication
Xiangyu Chen, Jixiang Luo, Yuankai Fan +3
Under the AI Flow framework, communication is shifting from transmitting fidelity-oriented information flows toward delivering task-oriented and perception-oriented token flows acr…
Rethinking Reward Signals in Video GRPO: When Scores Become Targets
Rui Li, Yuanzhi Liang, Ziqi Ni +3
The paper proposes TaRoS, a framework that redesigns reward signals for video generation using GRPO to avoid reward hacking and saturation, improving visual fidelity, motion cohere…