5 citations · 12 across the 25 of their papers we have counts for
1 paper · 1 filter
Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed
Multimodal large language models (MLLMs) increasingly answer questions whose correctness depends on visual, textual, temporal, acoustic, document, chart, or embodied evidence. Thei…