1 paper
Shuangquan Lyu, Jian Mao, Yue Ma
Diffusion-based text-to-video models are increasingly capable, but mask-based editing over hundreds of frames remains challenging: naïve long-video generation suffers from memory b…