14 citations · 23 across the 3 of their papers we have counts for
3 papers
cs.SD2023★ 9 cited
AudioToken: Adaptation of Text-Conditioned Diffusion Models for Audio-to-Image Generation
Guy Yariv, Itai Gat, Lior Wolf +2
In recent years, image generation has shown a great leap in performance, where diffusion models play a central role. Although generating high-quality images, such models are mainly…
cs.CV2022★ 14 cited
Zero-Shot Video Captioning with Evolving Pseudo-Tokens
Yoad Tewel, Yoav Shalev, Roy Nadler +2
We introduce a zero-shot video captioning method that employs two frozen networks: the GPT-2 language model and the CLIP image-text matching model. The matching score is used to st…
cs.LG2021
Latent Space Explanation by Intervention
Itai Gat, Guy Lorberbom, Idan Schwartz +1
The success of deep neural nets heavily relies on their ability to encode complex relations between their input and their output. While this property serves to fit the training dat…