1 paper
Adriano Fragomeni, Michael Wray, Dima Damen
In this paper, we re-examine the task of cross-modal clip-sentence retrieval, where the clip is part of a longer untrimmed video. When the clip is short or visually ambiguous, know…