1 paper · 1 filter
Peiyuan Zhang, Kaichen Zhang, Bo Li +7
Video sequences offer valuable temporal information, but existing large multimodal models (LMMs) fall short in understanding extremely long videos. Many works address this by reduc…