14 citations · 14 across the 3 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024
Detection-Driven Object Count Optimization for Text-to-Image Diffusion Models
Oz Zafar, Yuval Cohen, Lior Wolf +1
Accurately controlling object count in text-to-image generation remains a key challenge. Supervised methods often fail, as training data rarely covers all count variations. Methods…
cs.CV2022★ 14 cited
Zero-Shot Video Captioning with Evolving Pseudo-Tokens
Yoad Tewel, Yoav Shalev, Roy Nadler +2
We introduce a zero-shot video captioning method that employs two frozen networks: the GPT-2 language model and the CLIP image-text matching model. The matching score is used to st…