From the 1 of 4 linked papers with an AI index.
4 papers
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
Xinhao Li, Yuhan Zhu, Xiangyu Zeng +24
VideoChat3 is a fully open, 4B-parameter video-centric multimodal large language model that combines an efficient Inflated 3D Vision Transformer and adaptive frame resolution with…
CoMo: Learning Continuous Latent Motion from Internet Videos for Scalable Robot Learning
Jiange Yang, Yansong Shi, Haoyi Zhu +6
Unsupervised learning of latent motion from Internet videos is crucial for robot learning. Existing discrete methods generally mitigate the shortcut learning caused by extracting e…
RIVER: A Real-Time Interaction Benchmark for Video LLMs
Yansong Shi, Qingsong Zhao, Tianxiang Jiang +3
The rapid advancement of multimodal large language models has demonstrated impressive capabilities, yet nearly all operate in an offline paradigm, hindering real-time interactivity…
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
Xiangyu Zeng, Kunchang Li, Chenting Wang +10
Multimodal Large Language Models (MLLMs) have demonstrated impressive performance in short video understanding. However, understanding long-form videos still remains challenging fo…