17 citations · 18 across the 3 of their papers we have counts for
3 papers
cs.CV2024★ 17 cited
Lumiere: A Space-Time Diffusion Model for Video Generation
Omer Bar-Tal, Hila Chefer, Omer Tov +14
We introduce Lumiere -- a text-to-video diffusion model designed for synthesizing videos that portray realistic, diverse and coherent motion -- a pivotal challenge in video synthes…
cs.CV2024★ 1 cited
Inflation with Diffusion: Efficient Temporal Adaptation for Text-to-Video Super-Resolution
Xin Yuan, Jinoo Baek, Keyang Xu +2
We propose an efficient diffusion-based text-to-video super-resolution (SR) tuning approach that leverages the readily learned capacity of pixel level image diffusion model to capt…
cs.CV2023
Teaching CLIP to Count to Ten
Roni Paiss, Ariel Ephrat, Omer Tov +4
Large vision-language models (VLMs), such as CLIP, learn rich joint image-text representations, facilitating advances in numerous downstream tasks, including zero-shot classificati…