1 paper · 1 filter
Shihao Han, Hao Yang, Xinting Hu +3
Scaling Diffusion Transformers to generate high-resolution, long videos is constrained by the quadratic cost of self-attention, and existing sparse attention methods degrade under…