1 paper
Yunfeng Wu, Hongying Cheng, Zihao He +1
Transformer-based video diffusion models rely on 3D attention over spatial and temporal tokens, which incurs quadratic time and memory complexity and makes end-to-end training for…