1 paper
Dingxin Cheng, Mingda Li, Jingyu Liu +5
Recently, integrating visual foundation models into large language models (LLMs) to form video understanding systems has attracted widespread attention. Most of the existing models…