1 paper · 1 filter
Yuwei Fang, Willi Menapace, Aliaksandr Siarohin +5
Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datas…