1 paper
Ji Qi, Kaixuan Ji, Jifan Yu +4
Building models that comprehends videos and responds specific user instructions is a practical and challenging topic, as it requires mastery of both vision understanding and knowle…