14 citations · 14 across the 1 of their papers we have counts for
1 paper · 1 filter
Yoad Tewel, Yoav Shalev, Roy Nadler +2
We introduce a zero-shot video captioning method that employs two frozen networks: the GPT-2 language model and the CLIP image-text matching model. The matching score is used to st…