1 paper · 1 filter
Ziwei Zhou, Rui Wang, Zuxuan Wu +1
Recent Multimodal Large Language Models (MLLMs) achieve promising performance on visual and audio benchmarks independently. However, the ability of these models to process cross-mo…