24 papers
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification
Haopeng Li, Yitong Li, Junsong Chen +8
Diffusion transformers are essential for high-fidelity video generation, but long token sequences make attention a dominant inference bottleneck. Training-free dynamic sparse atten…
Trust Region Policy Distillation
Zhengpeng Xie, Li Lyna Zhang, Zeke Xie +1
Big goals are hard to achieve all at once; breaking them into small steps is wiser. We present Trust Region Policy Distillation (TOP-D), which transforms the notoriously unstable,…
MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model
Lichen Bai, Tianhao Zhang, Shitong Shao +14
As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but…
Optimizing Few-Step Generation with Adaptive Matching Distillation
Lichen Bai, Zikai Zhou, Shitong Shao +5
Distribution Matching Distillation (DMD) is a powerful acceleration paradigm, yet its stability is often compromised in Forbidden Zone, regions where the real teacher provides unre…
LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention
Shitong Shao, Zikai Zhou, Haopeng Li +4
Video editing has evolved toward In-Context Learning (ICL) paradigms, yet the resulting quadratic attention costs create a critical computational bottleneck. In this work, we propo…
Refining Multidimensional Video Reward Models via Disentangled Influence Functions
Muyao Wang, Zeke Xie, Hideki Nakayama
As Text-to-Video (T2V) generation models continue to evolve, the complexity of video evaluation necessitates a fine-grained assessment across various axes. To address this, recent…