1 citations · 2 across the 17 of their papers we have counts for
1 paper · 1 filter
Zhi Li, Yanan Wang, Hao Niu +2
Multimodal large language models have recently achieved remarkable progress in video question answering (VideoQA) by jointly processing visual, textual, and audio information. Howe…