4 papers · 1 filter
An LMM for Efficient Video Understanding via Reinforced Compression of Video Cubes
Ji Qi, Yuan Yao, Yushi Bai +4
Large Multimodal Models (LMMs) uniformly perceive video frames, creating computational inefficiency for videos with inherently varying temporal information density. This paper pres…
CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning
Ji Qi, Ming Ding, Weihan Wang +8
Vision-Language Models (VLMs) have demonstrated their broad effectiveness thanks to extensive training in aligning visual instructions to responses. However, such training of concl…
LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models
Shangqing Tu, Yucheng Wang, Daniel Zhang-Li +8
Existing Large Vision-Language Models (LVLMs) can process inputs with context lengths up to 128k visual and text tokens, yet they struggle to generate coherent outputs beyond 1,000…
VidCoM: Fast Video Comprehension through Large Language Models with Multimodal Tools
Ji Qi, Kaixuan Ji, Jifan Yu +4
Building models that comprehends videos and responds specific user instructions is a practical and challenging topic, as it requires mastery of both vision understanding and knowle…