activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

Hongbo Liu, Peixian Chen, Sihan Liu +12

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored.…

cs.CV2026

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs

Ye Wang, Hongjun Wang, Hao Fang +7

Unified multimodal models (UMMs) aim to integrate visual understanding and generation within a single architecture, but architectural unification alone does not ensure semantic con…

cs.CV2026

Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision

Zhixiang Wei, Yi Li, Zhehan Kan +38

Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, lea…

cs.CV2025

ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts

Yuying Ge, Yixiao Ge, Chen Li +15

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…

cs.CV2025

Beyond Intermediate States: Explaining Visual Redundancy through Language

Dingchen Yang, Bowen Cao, Anran Zhang +3

Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational b…

cs.CV2024

Video-Language Alignment via Spatio-Temporal Graph Transformer

Shi-Xue Zhang, Hongfa Wang, Xiaobin Zhu +5

Video-language alignment is a crucial multi-modal task that benefits various downstream applications, e.g., video-text retrieval and video question answering. Existing methods eith…