activity
20162023
most citedLearning What and Where to Draw

210 citations · 356 across the 21 of their papers we have counts for

collaborators

21 papers

eess.AS2023

Zero-shot audio captioning with audio-language model guidance and audio context keywords

Leonard Salewski, Stefan Fauth, A. Sophia Koepke +1

Zero-shot audio captioning aims at automatically generating descriptive textual captions for audio content without prior training for this task. Different from speech recognition w…

cs.CV2023

Zero-shot Translation of Attention Patterns in VQA Models to Natural Language

Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch +1

Converting a model's internals to text can yield human-understandable insights about the model. Inspired by the recent success of training-free approaches for image captioning, we…

cs.CV20231 cited

Transitivity Recovering Decompositions: Interpretable and Robust Fine-Grained Relationships

Abhra Chaudhuri, Massimiliano Mancini, Zeynep Akata +1

Recent advances in fine-grained representation learning leverage local-to-global (emergent) relationships for achieving state-of-the-art results. The relational representations rel…

cs.CV2023

Video-adverb retrieval with compositional adverb-action embeddings

Thomas Hummel, Otniel-Bogdan Mercea, A. Sophia Koepke +1

Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice…

cs.CV2023

Text-to-feature diffusion for audio-visual few-shot learning

Otniel-Bogdan Mercea, Thomas Hummel, A. Sophia Koepke +1

Training deep learning models for video classification from audio-visual data commonly requires immense amounts of labeled training data collected via a costly process. A challengi…

cs.CV2023

PDiscoNet: Semantically consistent part discovery for fine-grained recognition

Robert van der Klis, Stephan Alaniz, Massimiliano Mancini +4

Fine-grained classification often requires recognizing specific object parts, such as beak shape and wing patterns for birds. Encouraging a fine-grained classification model to fir…