3 citations · 5 across the 5 of their papers we have counts for
5 papers · 1 filter
Learning text-to-video retrieval from image captioning
Lucas Ventura, Cordelia Schmid, Gül Varol
We describe a protocol to study text-to-video retrieval training with unlabeled videos, where we assume (i) no access to labels for any videos, i.e., no access to the set of ground…
AutoAD III: The Prequel -- Back to the Pixels
Tengda Han, Max Bain, Arsha Nagrani +3
Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently, vi…
AutoAD II: The Sequel -- Who, When, and What in Movie Audio Description
Tengda Han, Max Bain, Arsha Nagrani +3
Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presen…
AutoAD: Movie Description in Context
Tengda Han, Max Bain, Arsha Nagrani +3
The objective of this paper is an automatic Audio Description (AD) model that ingests movies and outputs AD in text form. Generating high-quality movie AD is challenging due to the…
Automatic dense annotation of large-vocabulary sign language videos
Liliane Momeni, Hannah Bull, K R Prajwal +3
Recently, sign language researchers have turned to sign language interpreted TV broadcasts, comprising (i) a video of continuous signing and (ii) subtitles corresponding to the aud…