4 citations · 8 across the 3 of their papers we have counts for
4 papers · 1 filter
AutoAD III: The Prequel -- Back to the Pixels
Tengda Han, Max Bain, Arsha Nagrani +3
Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently, vi…
AutoAD II: The Sequel -- Who, When, and What in Movie Audio Description
Tengda Han, Max Bain, Arsha Nagrani +3
Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presen…
Balancing the Picture: Debiasing Vision-Language Datasets with Synthetic Contrast Sets
Brandon Smith, Miguel Farinha, Siobhan Mackenzie Hall +3
Vision-language models are growing in popularity and public visibility to generate, edit, and caption images at scale; but their outputs can perpetuate and amplify societal biases…
AutoAD: Movie Description in Context
Tengda Han, Max Bain, Arsha Nagrani +3
The objective of this paper is an automatic Audio Description (AD) model that ingests movies and outputs AD in text form. Generating high-quality movie AD is challenging due to the…