3 papers
cs.AI2026
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
Caorui Li, Yu Chen, Yiyan Ji +40
Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively eva…
cs.LG2025
Towards a Unified Representation Evaluation Framework Beyond Downstream Tasks
Christos Plachouras, Julien Guinot, George Fazekas +3
Downstream probing has been the dominant method for evaluating model representations, an important process given the increasing prominence of self-supervised learning and foundatio…
cs.SD2025
Learning Music Audio Representations With Limited Data
Christos Plachouras, Emmanouil Benetos, Johan Pauwels
Large deep-learning models for music, including those focused on learning general-purpose music audio representations, are often assumed to require substantial training data to ach…