5 citations · 6 across the 5 of their papers we have counts for
4 papers · 1 filter
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
Zhixiang Wei, Yi Li, Zhehan Kan +38
Despite the significant advancements represented by Vision-Language Models (VLMs), current architectures often exhibit limitations in retaining fine-grained visual information, lea…
ARC-Hunyuan-Video-7B: Structured Video Comprehension of Real-World Shorts
Yuying Ge, Yixiao Ge, Chen Li +15
Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal…
Beyond Intermediate States: Explaining Visual Redundancy through Language
Dingchen Yang, Bowen Cao, Anran Zhang +3
Multi-modal Large Langue Models (MLLMs) often process thousands of visual tokens, which consume a significant portion of the context window and impose a substantial computational b…
Adaptive Perception Transformer for Temporal Action Localization
Yizheng Ouyang, Tianjin Zhang, Weibo Gu +1
Temporal action localization aims to predict the boundary and category of each action instance in untrimmed long videos. Most of previous methods based on anchors or proposals negl…