2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Changli Tang, Qinfan Xiao, Ke Mei +3
While embeddings from multimodal large language models (LLMs) excel as general-purpose representations, their application to dynamic modalities like audio and video remains underex…