1 paper
Yanbo Ding, Yijia Fan, Caihua Shan +8
Diffusion Transformers (DiTs) have become the dominant paradigm for high-fidelity video generation, yet their ability to perform high-level semantic planning remains limited. While…