75 citations · 75 across the 1 of their papers we have counts for
1 paper
Simon Ging, Mohammadreza Zolfaghari, Hamed Pirsiavash +1
Many real-world video-text tasks involve different levels of granularity, such as frames and words, clip and sentences or videos and paragraphs, each with distinct semantics. In th…