19 citations · 39 across the 8 of their papers we have counts for
1 paper · 1 filter
Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu +1
Large language models (LLMs) are growingly extended to process multimodal data such as text and video simultaneously. Their remarkable performance in understanding what is shown in…