6 papers · 1 filter
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios
Hongbo Liu, Peixian Chen, Sihan Liu +12
Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored.…
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs
Ye Wang, Hongjun Wang, Hao Fang +7
Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic con…
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Zhixiang Wei, Yi Li, Zhehan Kan +38
Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, lea…
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Yuying Ge, Yixiao Ge, Chen Li +15
Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…
Beyond Intermediate States: Explaining Visual Redundancy through Language
Dingchen Yang, Bowen Cao, Anran Zhang +3
Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational b…
Video-Language Alignment via Spatio-Temporal Graph Transformer
Shi-Xue Zhang, Hongfa Wang, Xiaobin Zhu +5
Video-language alignment is a crucial multi-modal task that benefits various downstream applications, e.g., video-text retrieval and video question answering. Existing methods eith…