1 paper
Qilang Ye, Wei Zeng, Meng Liu +4
Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, whic…