6 papers · 1 filter
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
Qiankun Ma, Ziyao Zhang, Haofei Wang +3
Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead an…
UniAPO: Unified Multimodal Automated Prompt Optimization
Qipeng Zhu, Yanzhe Chen, Huasong Zhong +5
Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, dem…
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
Liuhan Chen, Xiaodong Cun, Xiaoyu Li +5
Video inbetweening aims to synthesize intermediate video sequences conditioned on the given start and end frames. Current state-of-the-art methods primarily extend large-scale pre-…
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
Jingwei Peng, Jiehao Chen, Mateo Alejandro Rojas +1
Complex Visual Question Answering (Complex VQA) tasks, which demand sophisticated multi-modal reasoning and external knowledge integration, present significant challenges for exist…
Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
Lei Tan, Weihao Li, Pingyang Dai +3
In the realm of Text-Based Person Search (TBPS), mainstream methods aim to explore more efficient interaction frameworks between text descriptions and visual data. However, recent…
PartFormer: Awakening Latent Diverse Representation from Vision Transformer for Object Re-Identification
Lei Tan, Pingyang Dai, Jie Chen +3
Extracting robust feature representation is critical for object re-identification to accurately identify objects across non-overlapping cameras. Although having a strong representa…