1 paper
Malsha Ashani Mahawatta Dona, Beatriz Cabrero-Daniel, Yinan Yu +1
Large language models (LLMs) are growingly extended to process multimodal data such as text and video simultaneously. Their remarkable performance in understanding what is shown in…