1 paper · 2 filters
Yifan Xu, Chao Zhang, Ruifei Ma +4
The new era has witnessed a remarkable capability to extend Vision-Language Models (VLMs) for tackling tasks of video understanding. While current VLMs excel at event- or story-lev…