1 paper · 1 filter
Zhiyuan Wu, Shuai Wang, Li Chen +5
Video diffusion models (VDMs) perform attention computation over the 3D spatio-temporal domain. Compared to large language models (LLMs) processing 1D sequences, their memory consu…