1 citations · 2 across the 4 of their papers we have counts for
1 paper · 1 filter
Yan Wang, Yawen Zeng, Jingsheng Zheng +3
Multimodal large language models (MLLMs) are flourishing, but mainly focus on images with less attention than videos, especially in sub-fields such as prompt engineering, video cha…