activity
20182026
most citedAudio Retrieval with Natural Language Queries: A Benchmark Study

77 citations · 123 across the 18 of their papers we have counts for

collaborators
Showing cs.CVShow all

16 papers · 1 filter

cs.CV2026

LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training

Andreas Hochlehnert, Marianna Nezhurina, Mehdi Cherti +9

We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download…

cs.CV2026

DF-MoE: Generalizable Deepfake Detection via Multimodal Sparse Mixture-of-Experts

Vlad Hondru, Florinel Alin Croitoru, Iuliana Georgescu +2

Audio-visual deepfake detection is an actively studied topic, where one of the main challenges is to develop detectors able to generalize across deepfake generation methods. We con…

cs.CV2026

Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale

A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1

The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same represent…

cs.CV2026

It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models

Anne Harrington, A. Sophia Koepke, Shyamgopal Karthik +2

Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prompt. Previous work has attempted…

cs.CV2024

Audio-Visual Generalized Zero-Shot Learning using Pre-Trained Large Multi-Modal Models

David Kurzendörfer, Otniel-Bogdan Mercea, A. Sophia Koepke +1

Audio-visual zero-shot learning methods commonly build on features extracted from pre-trained models, e.g. video or audio classification models. However, existing benchmarks predat…

cs.CV2023

Zero-shot Translation of Attention Patterns in VQA Models to Natural Language

Leonard Salewski, A. Sophia Koepke, Hendrik P. A. Lensch +1

Converting a model's internals to text can yield human-understandable insights about the model. Inspired by the recent success of training-free approaches for image captioning, we…