1 paper
Hangliang Ding, Dacheng Li, Runlong Su +4
Despite the promise of synthesizing high-fidelity videos, Diffusion Transformers (DiTs) with 3D full attention suffer from expensive inference due to the complexity of attention co…