most citedExposing and Mitigating Spurious Correlations for Cross-Modal Retrieval

4 citations · 6 across the 7 of their papers we have counts for

collaborators

9 papers

cs.CV2024

Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models

David Kurzendörfer, Otniel-Bogdan Mercea, A. Sophia Koepke +1

Audio-visual zero-shot learning methods commonly build on features extracted from pre-trained models, e.g. video or audio classification models. However, existing benchmarks predat…

eess.AS2024

A SOUND APPROACH: Using Large Language Models to generate audio descriptions for egocentric text-audio retrieval

Andreea-Maria Oncescu, João F. Henriques, Andrew Zisserman +2

Video databases from the internet are a valuable source of text-audio retrieval datasets. However, given that sound and vision streams represent different "views" of the data, trea…

eess.AS2023

Zero-shot audio captioning with audio-language model guidance and audio context keywords

Leonard Salewski, Stefan Fauth, A. Sophia Koepke +1

Zero-shot audio captioning aims at automatically generating descriptive textual captions for audio content without prior training for this task. Different from speech recognition w…

cs.CV2023

Zero-shot Translation of Attention Patterns in VQA Models to Natural Language

Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch +1

Converting a model's internals to text can yield human-understandable insights about the model. Inspired by the recent success of training-free approaches for image captioning, we…

cs.CV2023

Video-adverb retrieval with compositional adverb-action embeddings

Thomas Hummel, Otniel-Bogdan Mercea, A. Sophia Koepke +1

Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice…

cs.CV2023

Text-to-feature diffusion for audio-visual few-shot learning

Otniel-Bogdan Mercea, Thomas Hummel, A. Sophia Koepke +1

Training deep learning models for video classification from audio-visual data commonly requires immense amounts of labeled training data collected via a costly process. A challengi…