3 papers
cs.MM2026
VGGSounder: Audio-Visual Evaluations for Foundation Models
Daniil Zverev, Thaddäus Wiedemer, Ameya Prabhu +3
The emergence of audio-visual foundation models underscores the importance of reliably assessing their multi-modal understanding. The VGGSound dataset is commonly used as a benchma…
cs.CV2026
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
A. Sophia Koepke, Daniil Zverev, Shiry Ginosar +1
The Platonic Representation Hypothesis suggests that neural networks trained on different modalities (e.g., text and images) align and eventually converge toward the same represent…
cs.LG2025
On the Dangers of Bootstrapping Generation for Continual Learning and Beyond
Daniil Zverev, A. Sophia Koepke, Joao F. Henriques
The use of synthetically generated data for training models is becoming a common practice. While generated data can augment the training data, repeated training on synthetic data r…