5 papers
DUET: A Diversity-Quality Duet of Distillation Experts for Two-Step Video Generation
Zian Li, Litong Gong, Borui Liao +6
Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment. Few-step distillation allev…
Edit-GRPO: A Locality-Preserving Policy Optimization Framework for Image Editing
Shaodong Xu, Zexian Li, Zhendong Wang +5
A fundamental challenge in image editing lies in preserving spatial locality: edits should improve targeted content without inadvertently altering surrounding regions. However, mos…
Beyond Point-Wise Matching: Structural Representation Alignment for Accelerating Diffusion Transformers
Shaodong Xu, Zhendong Wang, Litong Gong +4
Recent advances in Diffusion Transformers (DiTs) demonstrate that aligning noisy latent states with well-trained semantic features-as pioneered by Representation Alignment (REPA)-c…
AdvDMD: Adversarial Reward Meets DMD For High-Quality Few-Step Generation
Xu Wang, Zexian Li, Litong Gong +2
Diffusion models offer superior generation quality at the expense of extensive sampling steps. Distillation methods, with Distribution Matching Distillation (DMD) as a popular exam…
RISE-T2V: Rephrasing and Injecting Semantics with LLM for Expansive Text-to-Video Generation
Xiangjun Zhang, Litong Gong, Yinglin Zheng +6
Most text-to-video(T2V) diffusion models depend on pre-trained text encoders for semantic alignment, yet they often fail to maintain video quality when provided with concise prompt…