1 citations · 1 across the 25 of their papers we have counts for
1 paper · 1 filter
Yifan Dai, Zhenhua Wu, Bohan Zeng +18
Joint audio-visual reasoning is essential for omnimodal understanding, yet current multimodal large language models (MLLMs) still struggle when reasoning requires fine-grained evid…