17 papers
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
Lichen Bai, Tianhao Zhang, Shitong Shao +14
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but…
LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention
Shitong Shao, Zikai Zhou, Haopeng Li +4
Video editing has evolved toward In-Context Learning (ICL) paradigms, yet the resulting quadratic attention costs create a critical computational bottleneck. In this work, we propo…
Exploring Data-Free LoRA Transferability for Video Diffusion Models
Yuchen Wang, Wenliang Zhong, Lichen Bai +6
Video diffusion models leveraging step distillation or causal distillation have achieved remarkable performance. However, adapting existing LoRAs to these variants remains a critic…
Reflective Flow Sampling Enhancement
Zikai Zhou, Muyao Wang, Shitong Shao +4
The growing demand for text-to-image generation has led to rapid advances in generative modeling. Recently, text-to-image diffusion models trained with flow matching algorithms, su…
Efficient Video Diffusion Models: Advancements and Challenges
Shitong Shao, Lichen Bai, Pengfei Wan +2
Video diffusion models have rapidly become the dominant paradigm for high-fidelity generative video synthesis, but their practical deployment remains constrained by severe inferenc…
CRAFT: Aligning Diffusion Models with Fine-Tuning Is Easier Than You Think
Zening Sun, Zhengpeng Xie, Lichen Bai +3
Aligning Diffusion models has achieved remarkable breakthroughs in generating high-quality, human preference-aligned images. Existing techniques, such as supervised fine-tuning (SF…