1 paper
Yuxi Liu, Haoyu Li, Zekun Zhang +12
Video diffusion transformers are expensive because attention dominates long spatiotemporal token sequences. We identify the \emph{high-sparsity trap}: at extreme attention sparsity…