2 papers
cs.CV2026
UFVideo: Towards Unified Fine-Grained Video Cooperative Understanding with Large Language Models
Hewen Pan, Cong Wei, Dashuang Liang +8
With the advancement of multi-modal Large Language Models (LLMs), Video LLMs have been further developed to perform on holistic and specialized video understanding. However, existi…
cs.CV2025
DisTime: Distribution-based Time Representation for Video Large Language Models
Yingsen Zeng, Zepeng Huang, Yujie Zhong +4
Despite advances in general video understanding, Video Large Language Models (Video-LLMs) face challenges in precise temporal localization due to discrete time representations and…