25 citations · 33 across the 8 of their papers we have counts for
Showing 2024Show all
2 papers · 1 filter
cs.CV2024
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
Junyu Xie, Tengda Han, Max Bain +4
Our objective is to generate Audio Descriptions (ADs) for both movies and TV series in a training-free manner. We use the power of off-the-shelf Visual-Language Models (VLMs) and L…
cs.CV2024
AutoAD III: The Prequel -- Back to the Pixels
Tengda Han, Max Bain, Arsha Nagrani +3
Generating Audio Description (AD) for movies is a challenging task that requires fine-grained visual understanding and an awareness of the characters and their names. Currently, vi…