1 citations · 1 across the 3 of their papers we have counts for
3 papers
Zero-shot audio captioning with audio-language model guidance and audio context keywords
Leonard Salewski, Stefan Fauth, A. Sophia Koepke +1
Zero-shot audio captioning aims at automatically generating descriptive textual captions for audio content without prior training for this task. Different from speech recognition w…
Zero-shot Translation of Attention Patterns in VQA Models to Natural Language
Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch +1
Converting a model's internals to text can yield human-understandable insights about the model. Inspired by the recent success of training-free approaches for image captioning, we…
Diverse Video Captioning by Adaptive Spatio-temporal Attention
Zohreh Ghaderi, Leonard Salewski, Hendrik P. A. Lensch
To generate proper captions for videos, the inference needs to identify relevant concepts and pay attention to the spatial relationships between them as well as to the temporal dev…