4 papers
VGGSounder: Audio-Visual Evaluations for Foundation Models
Daniil Zverev, Thaddäus Wiedemer, Ameya Prabhu +3
The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchma…
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1
The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same represent…
It's Never Too Late: Noise Optimization for Collapse Recovery in Trained Diffusion Models
Anne Harrington, A. Sophia Koepke, Shyamgopal Karthik +2
Contemporary text-to-image models exhibit a surprising degree of mode collapse, as can be seen when sampling several images given the same text prompt. Previous work has attempted…
On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
Daniil Zverev, A. Sophia Koepke, Joao F. Henriques
The use of synthetically generated data for training models is becoming a common practice. While generated data can augment the training data, repeated training on synthetic data r…