1 paper
Kyobin Choo, Youngmin Kim, Hyunkyung Han +4
Video diffusion transformers (DiTs) generate high-fidelity and temporally coherent videos, yet motion control remains implicit, primarily relying on text prompts. As a result, achi…