6 papers · 1 filter
WorldOlympiad: Can Your World Model Survive a Triathlon?
Yuke Zhao, Wangbo Zhao, Weijie Wang +8
We introduce WorldOlympiad, a benchmark for diagnosing video-based world models across physical faithfulness, geometric consistency, and interaction fidelity. While existing benchm…
BIFE: Better Interaction, Fewer Errors for Minute-Long Video Generation
Zeyu Zhang, Jinyuan Mao, Shuning Chang +5
Long video generation is a critical step toward building realistic world models, requiring both high visual fidelity and long-range interaction consistency. Recent autoregressive d…
Towards Error-Free Long Video Generation
Shuning Chang, Weihua Chen, Jiasheng Tang +8
Recent advances in video generation have made minute-level synthesis possible; however, generating long videos remains challenging due to error accumulation, attribute drift, and t…
Few-Step Distillation for Text-to-Image Generation: A Practical Guide
Yifan Pu, Yizeng Han, Zhiwei Tang +4
Diffusion distillation has dramatically accelerated class-conditional image synthesis, but its applicability to open-ended text-to-image (T2I) generation is still unclear. We prese…
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
Wangbo Zhao, Yizeng Han, Zhiwei Tang +7
Diffusion Transformers (DiTs) excel at visual generation yet remain hampered by slow sampling. Existing training-free accelerators - step reduction, feature caching, and sparse att…
SparseDiT: Token Sparsification for Efficient Diffusion Transformer
Shuning Chang, Pichao Wang, Jiasheng Tang +2
Diffusion Transformers (DiT) are renowned for their impressive generative performance; however, they are significantly constrained by considerable computational costs due to the qu…