1 paper · 1 filter
Tianhong Zhou, Mingyang Han, Boyu Li +8
Audio-visual feature extraction is a fundamental component of multimodal understanding and generation tasks. However, existing evaluation protocols for feature extraction models ex…