4 papers
TimeChat-Online: 80% Visual Tokens are Naturally Redundant in Streaming Videos
Linli Yao, Yicheng Li, Yuancheng Wei +11
The rapid growth of online video platforms, particularly live streaming services, has created an urgent need for real-time video understanding systems. These systems must process c…
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
Shicheng Li, Lei Li, Kun Ouyang +7
Video Large Language Models (Video LLMs) have achieved significant success by adopting the paradigm of large-scale pre-training followed by supervised fine-tuning (SFT). However, e…
PunchBench: Benchmarking MLLMs in Multimodal Punchline Comprehension
Kun Ouyang, Yuanxin Liu, Shicheng Li +5
Multimodal punchlines, which involve humor or sarcasm conveyed in image-caption pairs, are a popular way of communication on online multimedia platforms. With the rapid development…
DCA: Diversified Co-Attention towards Informative Live Video Commenting
Zhihan Zhang, Zhiyi Yin, Shuhuai Ren +2
We focus on the task of Automatic Live Video Commenting (ALVC), which aims to generate real-time video comments with both video frames and other viewers' comments as inputs. A majo…