4 citations · 8 across the 3 of their papers we have counts for
3 papers
cs.CV2023★ 1 cited
AutoAD II: The Sequel -- Who, When, and What in Movie Audio Description
Tengda Han, Max Bain, Arsha Nagrani +3
Audio Description (AD) is the task of generating descriptions of visual content, at suitable time intervals, for the benefit of visually impaired audiences. For movies, this presen…
cs.CV2023★ 4 cited
Balancing the Picture: Debiasing Vision-Language Datasets with Synthetic Contrast Sets
Brandon Smith, Miguel Farinha, Siobhan Mackenzie Hall +3
Vision-language models are growing in popularity and public visibility to generate, edit, and caption images at scale; but their outputs can perpetuate and amplify societal biases…
cs.CV2023★ 3 cited
AutoAD: Movie Description in Context
Tengda Han, Max Bain, Arsha Nagrani +3
The objective of this paper is an automatic Audio Description (AD) model that ingests movies and outputs AD in text form. Generating high-quality movie AD is challenging due to the…