1 citations · 1 across the 1 of their papers we have counts for
1 paper
Han Wang, Yanjie Wang, Yongjie Ye +2
Multi-modal Large Language Models (MLLMs) have demonstrated their ability to perceive objects in still images, but their application in video-related tasks, such as object tracking…