1 paper · 1 filter
Yanbo Ding, Yijia Fan, Caihua Shan +8
Diffusion Transformers (DiTs) have become the dominant paradigm for high-fidelity video generation, yet their ability to perform high-level semantic planning remains limited. While…