1 paper · 1 filter
Yong Ren, Chenxing Li, Le Xu +7
Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remain…