1 paper · 1 filter
Yuxuan Wang, Yueqian Wang, Pengfei Wu +4
Despite progress in multimodal large language models (MLLMs), the challenge of interpreting long-form videos in response to linguistic queries persists, largely due to the ineffici…