6 papers
ApET: Approximation-Error Guided Token Compression for Efficient VLMs
Qiankun Ma, Ziyao Zhang, Haofei Wang +3
Recent Vision-Language Models (VLMs) have demonstrated remarkable multimodal understanding capabilities, yet the redundant visual tokens incur prohibitive computational overhead an…
UniAPO: Unified Multimodal Automated Prompt Optimization
Qipeng Zhu, Yanzhe Chen, Huasong Zhong +5
Prompting is fundamental to unlocking the full potential of large language models. To automate and enhance this process, automatic prompt optimization (APO) has been developed, dem…
EF-VI: Enhancing End-Frame Injection for Video Inbetweening
Liuhan Chen, Xiaodong Cun, Xiaoyu Li +5
Video inbetweening aims to synthesize intermediate video sequences conditioned on the given start and end frames. Current state-of-the-art methods primarily extend large-scale pre-…
MV-CoRe: Multimodal Visual-Conceptual Reasoning for Complex Visual Question Answering
Jingwei Peng, Jiehao Chen, Mateo Alejandro Rojas +1
Complex Visual Question Answering (Complex VQA) tasks, which demand sophisticated multi-modal reasoning and external knowledge integration, present significant challenges for exist…
TRiMM: Transformer-Based Rich Motion Matching for Real-Time multi-modal Interaction in Digital Humans
Yueqian Guo, Tianzhao Li, Xin Lyu +7
Large Language Model (LLM)-driven digital humans have sparked a series of recent studies on co-speech gesture generation systems. However, existing approaches struggle with real-ti…
Knowing Where to Focus: Attention-Guided Alignment for Text-based Person Search
Lei Tan, Weihao Li, Pingyang Dai +3
In the realm of Text-Based Person Search (TBPS), mainstream methods aim to explore more efficient interaction frameworks between text descriptions and visual data. However, recent…